Generative transition
Spawn sampling and the one-decision transition.
Duckietown._uniform_tile_index — Method
_uniform_tile_index(rng, n) -> Intnp_random.integers(0, n) for the random start-tile choice. A NumpyPCG64 takes the reference (buffered-Lemire) path; any other RNG uses the native uniform draw.
Duckietown.drivable_tiles — Method
drivable_tiles(map) -> Vector{NTuple{2,Int}}Zero-based (i, j) coordinates of the drivable tiles, in the reference row-major insertion order of Simulator._interpret_map.
Duckietown.initial_duckie — Method
initial_duckie(map, cfg) -> DuckieStateThe dynamic duckie the controller injects (duck_controller._inject_duck + Simulator.interpret_object + DuckieObj.__init__ with domain_rand=false): mesh geometry from the tile-frame descriptor, vel = 0.02, pedestrian_wait_time = 8, pedestrian_active = false, time = 0, start = center = pos, heading = heading_vec(angle).
walk_distance is cfg.walk_distance because DuckController.__init__ stamps it over the simulator's default (road_tile_size) — the FJ3.5 finding.
Duckietown.initial_map — Method
initial_map(cfg) -> RoadMapThe experiment map with the controller's injected static stop sign (duck_controller._inject_stop, tile-frame descriptor), when the config asks for one. Duckies are dynamic state, not map objects.
Duckietown.position_in_bounds_xz — Method
position_in_bounds_xz(pos, bounds) -> Boolenv_wrapper.position_in_bounds_xz with bounds = (xmin, xmax, zmin, zmax).
Duckietown.route_circulation_score — Method
route_circulation_score(pos, angle, center) -> Float64env_wrapper.route_circulation_score: alignment of the ego heading with the clockwise tangent around center (x-z plane). Positive = clockwise. Used only to constrain the initial-state distribution; never part of the state.
Duckietown.sample_spawn_pose — Method
sample_spawn_pose(map, rng, i, j, accept_start_angle_deg) -> (pos, angle)Simulator reset spawn sampling on tile (i, j) (world coordinates, x/z within [i·ts, (i+1)·ts)). Returns (pos3, angle); on exhaustion of MAX_SPAWN_ATTEMPTS the reference fallback ([1.0, 0.0, 1.0], 1.0).
Duckietown.DuckieTransitionModel — Type
DuckieTransitionModelStatic parameters of the one-decision transition — everything DuckieMDPEnv/ContinuousDuckieMDPEnv reads besides the world state: action/state/reward/duck-controller/continuous configs, frame_skip, max_steps (physics ticks, compared against ego.step_count), and the optional goal_tile. Built from a DuckietownConfig so every parameter traces back to a single experiment YAML.
Duckietown.TerminationReason — Type
TerminationReasonExplicit single termination reason, in the reference resolution order (duck_collision > other_collision > timeout > offroad > goal > in_progress); an object collision is never conflated with going off-road.
Duckietown.TransitionResult — Type
TransitionResultRich output of one generative decision. sp is the successor world state; everything else is the audit trail the wrapper's info dict carries: solver-facing projections, the component-level reward, the event flags, the terminated/truncated split (timeout is truncation, not an absorbing state), the explicit reason, and the wheel commands actually applied (Float32, post-clip).
Duckietown._duck_collision — Method
_duck_collision(world) -> Boolenv_wrapper._duck_collision: SAT of the agent bounding box against every visible duckie's shifted corners.
Duckietown.is_terminated — Method
is_terminated(reason) -> Bool
is_truncated(reason) -> BoolThe TD-bootstrapping split: a genuine terminal (duck_collision, other_collision, offroad, goal) breaks the bootstrap; timeout is only the experiment horizon (truncation), not an absorbing physical state.
Duckietown.simulate_decision — Method
simulate_decision(model, s, action_id::Integer, rng) -> TransitionResult
simulate_decision(model, s, action::DuckieAction, rng) -> TransitionResultOne full macro-decision of the Duckietown MDP from world state s (DuckieMDPEnv.step for a discrete action_id in 0:6, ContinuousDuckieMDPEnv.step for a continuous [v_cmd, omega_cmd]).
s is never mutated: every step of the chain builds fresh state (branch-pure), so a planner may call this repeatedly from the same s with different actions. Stochasticity (the p_cross trigger draw) comes from the external rng — the MDP semantics are x' ~ T(.|x, a) with the noise supplied by the caller, not stored in the state.
Continuous actions are clipped to the reference Box ([0, v_fast] x [-w0, w0], Float32 like np.clip on the float32 command) and the steering penalty uses the pre-action curvature of s with the clipped omega_cmd (README design constraint #3).
Duckietown.termination_reason — Method
termination_reason(model, world) -> TerminationReasonThe single explicit reason for world, in the reference resolution order (DuckieMDPEnv.step). Every input is a function of the post-transition world state alone — simulator_done (Simulator._compute_done_reward: invalid pose, which already covers collisions and off-road, or the physics-tick horizon), the duckie SAT test, the full collision test, and the goal tile — so the same function serves the transition chain and POMDPs.isterminal; there is no second copy of this logic.