Reward & events

The shaped reward, the stop tracker, and the event flags.

Duckietown.EventFlags — Type
EventFlags

Discrete events that add a bonus or penalty on top of the dense reward (src/reward.py::EventFlags). passed_stop is informational only; it never enters the reward. timeout marks truncation (non-absorbing).

source
Duckietown.RewardBreakdown — Type
RewardBreakdown

Per-component reward decomposition (src/reward.py::RewardBreakdown) in the same field order as the Python dataclass: progress, lateral, heading, time, pedestrian, stagnation, stop_approach, steering, events, total.

source
Duckietown.RewardConfig — Type
RewardConfig

All reward coefficients and thresholds (src/reward.py::RewardConfig). Defaults mirror Python exactly. The four experiment YAMLs override a subset; anything absent from the YAML falls back to these Python defaults (this is the authoritative config hierarchy: YAML > source defaults, with defaults applied per missing key, exactly as RewardConfig(**config["reward"]) behaves).

Terms (evaluated on the post-transition state, except the continuous steering term which uses the pre-action kappa and the clipped ω_cmd; see docs/src/validation/FJ1_STATUS.md):

  • dense: progress = α_p·v·cos(phi), lateral = -α_d·d², heading = -α_φ·phi², time = -c_step
  • pedestrian: duck_yield if v < duck_yield_speed during a crossing else duck_unsafe
  • stagnation: unnecessary_stop if v < idle_speed and not crossing and not must-stop
  • stop approach: stop_approach_yield/stop_approach_unsafe within stop_approach_distance of an unmet stop (disabled at distance 0)
  • steering: -|straight_steer_penalty|·min(1,|ω|/max_steer_command)² on straight segments
  • events: collision/offroad/stopviolation/fullstop/goal coefficients
source
Duckietown.compute_reward — Function
compute_reward(state, events, cfg=RewardConfig(); action_omega=0.0, curvature=nothing) -> RewardBreakdown

Dense per-decision reward R(s, a, s') with per-component decomposition (src/reward.py::compute_reward, exact term order and arithmetic):

r = α_p·v·cos(phi) - α_d·d² - α_φ·phi² - c_step + r_event
  • pedestrian: during a crossing (duck ∈ {CROSSING_FAR, CROSSING_NEAR}), duck_yield if v < duck_yield_speed else duck_unsafe.
  • stagnation: unnecessary_stop if v < idle_speed and not crossing and not within the stop-hold zone of an unmet stop.
  • stop_approach: within stop_approach_distance of an unmet stop, stop_approach_yield if v < stop_approach_speed else stop_approach_unsafe (disabled at distance 0, preserving every older baseline).
  • steering: on straight geometry (|curvature| ≤ straight_curvature_threshold, curvature supplied), -|penalty|·min(1, |ω_cmd|/max_steer_command)². The curvature is the pre-action kappa from s_t and action_omega the clipped command — this keeps the term in R(s, a, s') form and must never read sp.kappa (locked FJ1 constraint 3).
  • events: linear combination of the discrete event flags with their coefficients, in Python's addition order.

All arithmetic is Float64, matching NumPy.

source
Duckietown.StopTracker — Type
StopTracker

Dwell-counter state of the stop-compliance mechanism (src/reward.py::StopTracker).

  • zone: distance (m) within which a slow ego may accumulate dwell.
  • speed: velocity (m/s) below which the ego counts as slow.
  • pass_distance: distance (m) below which the stop is considered passed; clamped to ≥ zone (Python: pass_distance = max(zone, pass_distance)).
  • hold_steps_required: dwell steps needed to set sigma_stop; clamped to ≥ 1 (experiments: 1 tabular, 3 SAC/TD3).

The dwell counter is a mutable memory cell, updated in place by update! exactly like the Python attribute mutations; rollouts must call it on a branched copy (see branch).

source
Duckietown.hold_progress — Method
hold_progress(tracker) -> Float64

Normalized dwell progress: 1.0 once stopped, else hold_steps / hold_steps_required clipped to [0, 1]. This feature keeps the process Markov (append-only 15th continuous feature).

source
Duckietown.update! — Function
update!(tracker, previous, current, previous_stop_id=nothing, current_stop_id=nothing) -> (sigma_stop, events)

Advance the stop memory over one macro-decision and return (sigma_stop, events) (src/reward.py::StopTracker.update, exact semantics).

Passed detection (awarded once; resets dwell):

  • with stop ids available (previous_stop_id !== nothing): passed iff the sign changed AND previous.d_stop ≤ pass_distance;
  • without ids: passed iff previous.d_stop ≤ pass_distance AND (current has no stop OR its distance jumped > 0.5 beyond the previous).

A passed stop sets passed_stop, and stop_violation iff sigma_stop was not yet set. A sign change without a pass resets the memory. Otherwise the dwell increments only while near (d_stop ≤ zone) AND slow (v < speed), must be consecutive (hold_steps resets on any non-qualifying step), and latches sigma_stop at hold_steps_required steps with full_stop raised exactly once.

Mirrors Python exactly, including mutating the tracker in place (the canonical dwell memory lives in the world state; call this on the branched state in rollouts).

source