Reward & events
The shaped reward, the stop tracker, and the event flags.
Duckietown.EventFlags — Type
EventFlagsDiscrete events that add a bonus or penalty on top of the dense reward (src/reward.py::EventFlags). passed_stop is informational only; it never enters the reward. timeout marks truncation (non-absorbing).
Duckietown.RewardBreakdown — Type
RewardBreakdownPer-component reward decomposition (src/reward.py::RewardBreakdown) in the same field order as the Python dataclass: progress, lateral, heading, time, pedestrian, stagnation, stop_approach, steering, events, total.
Duckietown.RewardConfig — Type
RewardConfigAll reward coefficients and thresholds (src/reward.py::RewardConfig). Defaults mirror Python exactly. The four experiment YAMLs override a subset; anything absent from the YAML falls back to these Python defaults (this is the authoritative config hierarchy: YAML > source defaults, with defaults applied per missing key, exactly as RewardConfig(**config["reward"]) behaves).
Terms (evaluated on the post-transition state, except the continuous steering term which uses the pre-action kappa and the clipped ω_cmd; see docs/src/validation/FJ1_STATUS.md):
- dense:
progress = α_p·v·cos(phi),lateral = -α_d·d²,heading = -α_φ·phi²,time = -c_step - pedestrian:
duck_yieldifv < duck_yield_speedduring a crossing elseduck_unsafe - stagnation:
unnecessary_stopifv < idle_speedand not crossing and not must-stop - stop approach:
stop_approach_yield/stop_approach_unsafewithinstop_approach_distanceof an unmet stop (disabled at distance 0) - steering:
-|straight_steer_penalty|·min(1,|ω|/max_steer_command)²on straight segments - events: collision/offroad/stopviolation/fullstop/goal coefficients
Duckietown.compute_reward — Function
compute_reward(state, events, cfg=RewardConfig(); action_omega=0.0, curvature=nothing) -> RewardBreakdownDense per-decision reward R(s, a, s') with per-component decomposition (src/reward.py::compute_reward, exact term order and arithmetic):
r = α_p·v·cos(phi) - α_d·d² - α_φ·phi² - c_step + r_eventpedestrian: during a crossing (duck ∈ {CROSSING_FAR, CROSSING_NEAR}),duck_yieldifv < duck_yield_speedelseduck_unsafe.stagnation:unnecessary_stopifv < idle_speedand not crossing and not within the stop-hold zone of an unmet stop.stop_approach: withinstop_approach_distanceof an unmet stop,stop_approach_yieldifv < stop_approach_speedelsestop_approach_unsafe(disabled at distance 0, preserving every older baseline).steering: on straight geometry (|curvature| ≤ straight_curvature_threshold,curvaturesupplied),-|penalty|·min(1, |ω_cmd|/max_steer_command)². The curvature is the pre-actionkappafroms_tandaction_omegathe clipped command — this keeps the term inR(s, a, s')form and must never readsp.kappa(locked FJ1 constraint 3).events: linear combination of the discrete event flags with their coefficients, in Python's addition order.
All arithmetic is Float64, matching NumPy.
Duckietown.StopTracker — Type
StopTrackerDwell-counter state of the stop-compliance mechanism (src/reward.py::StopTracker).
zone: distance (m) within which a slow ego may accumulate dwell.speed: velocity (m/s) below which the ego counts as slow.pass_distance: distance (m) below which the stop is considered passed; clamped to≥ zone(Python:pass_distance = max(zone, pass_distance)).hold_steps_required: dwell steps needed to setsigma_stop; clamped to≥ 1(experiments: 1 tabular, 3 SAC/TD3).
The dwell counter is a mutable memory cell, updated in place by update! exactly like the Python attribute mutations; rollouts must call it on a branched copy (see branch).
Duckietown.hold_progress — Method
hold_progress(tracker) -> Float64Normalized dwell progress: 1.0 once stopped, else hold_steps / hold_steps_required clipped to [0, 1]. This feature keeps the process Markov (append-only 15th continuous feature).
Duckietown.reset_tracker — Method
reset_tracker(tracker) -> StopTrackerFresh tracker with the same configuration and cleared dwell memory.
Duckietown.update! — Function
update!(tracker, previous, current, previous_stop_id=nothing, current_stop_id=nothing) -> (sigma_stop, events)Advance the stop memory over one macro-decision and return (sigma_stop, events) (src/reward.py::StopTracker.update, exact semantics).
Passed detection (awarded once; resets dwell):
- with stop ids available (
previous_stop_id !== nothing): passed iff the sign changed ANDprevious.d_stop ≤ pass_distance; - without ids: passed iff
previous.d_stop ≤ pass_distanceAND (currenthas no stop OR its distance jumped> 0.5beyond the previous).
A passed stop sets passed_stop, and stop_violation iff sigma_stop was not yet set. A sign change without a pass resets the memory. Otherwise the dwell increments only while near (d_stop ≤ zone) AND slow (v < speed), must be consecutive (hold_steps resets on any non-qualifying step), and latches sigma_stop at hold_steps_required steps with full_stop raised exactly once.
Mirrors Python exactly, including mutating the tracker in place (the canonical dwell memory lives in the world state; call this on the branched state in rollouts).