FJ3 — Native Dynamics, Observers, Transition, RNG
Last update: 2026-08-19. Status: FJ3.1–FJ3.8 done — FJ3 COMPLETE.
RNG-C claim, in the wording to use in the README/paper:
Bit-exact compatibility for the tested NumPy RNG streams and exercised generator paths; rare libm-dependent normal-sampling branches retain the documented cross-platform ≤1-ULP caveat.
(The ziggurat fast path — ~99.3% of standard_normal draws — is integer × table and therefore bit-exact; only the wedge/tail rejection branches call log/exp. Nothing in the MDP transition chain draws normals at all: the controller stream uses random_sample only.)
FJ3 ports the latent dynamics and the env-dependent state extraction that FJ2 could not pin as pure functions. Every substep is pinned against fixtures generated by the real reference stack in the ddm-ref conda env (Python 3.9, numpy 1.20.0, gym 0.23.1, duckietown-gym-daffy 6.1.34, duckietown-world-daffy 6.4.3) plus the duckduck/src modules.
Substep table
| Substep | Scope | Fixture / test | Status |
|---|---|---|---|
| FJ3.1 | Map reconstruction (small_loop tiles, Bézier lanes, closest_curve_point, get_lane_pos2, collision geometry, spawn sampling) | fj3_map.json / test_fj3_map.jl | done |
| FJ3.2 | DB18 nominal motor model + 0.15 s delayed dynamics, SE(2) chain, weird_from_cartesian | fj3_ego.json / test_fj3_ego.jl | done |
| FJ3.3 | Delay-buffer window semantics (trim, u0 fallback, cross-decision memory) | fj3_ego.json / test_fj3_delay.jl | done |
| FJ3.4 | DuckieObj.step walk / finish_walk dynamics (600-tick standalone scenarios) | fj3_duck.json / test_fj3_duck.jl | done |
| FJ3.5 | DuckController.before_step trigger semantics + full 137-decision rollout chain parity (ego ticks + duck ticks + trigger) | fj3_duck.json / test_fj3_duck.jl | done |
| FJ3.6 | State-extraction observers: get_raw_state, next_stop_candidate, classify_duck, tile_ahead, both _lane_frame variants, duck_relative_state, signed_curvature_ahead, get_continuous_state, encode | fj3_obs.json (334 rows) / test_fj3_obs.jl (18 994 assertions) | done |
| FJ3.7 | One-decision transition chain (simulate_decision → TransitionResult): before_step → wheels → frame_skip ticks → sim-done → extraction → StopTracker → collision/reason → terminated/truncated → reward; discrete AND continuous | fj37_transition.json (24 scenarios, 1250 decisions, generated by the REAL DuckieMDPEnv.step/ContinuousDuckieMDPEnv.step) / test_fj3_transition.jl (59 913 assertions) | done |
| FJ3.8 | Exact NumPy RNG stream identity: NumpyMT19937 (= np.random.RandomState) and NumpySeedSequence/NumpyPCG64 (= gym.utils.seeding.np_random) + the Generator methods the reset path calls | fj38_rng.json / test_fj3_rng.jl (4 714 assertions) | done — RNG-C exact |
FJ3.6 design
Fixture rows are latent-in → outputs-out: each row records the full latent input (ego pose/speed, duck center/heading/vel/active/visible, controller counters, sigma/fallback memory) and every observer output. The Julia test rebuilds a DuckieWorldState from the row, so extraction parity is pinned without accumulating dynamics drift (chain-level integration is FJ3.7). 300 rows from a driven 3-lap rollout + 34 synthetic rows covering: every next_stop_candidate filter branch (distance/behind/lateral/ orientation), all 5 DuckThreat classes, all 3 ego-relative tile classes, the detection gate, controller-counter crossing_available, the NotInLane fallback, and the off-map STRAIGHT fallback.
New state field: DuckieEgoState.speed — the simulator's self.speed = norm(cur_pos - prev_pos) / delta_time (displacement speed of the last tick). This is what get_raw_state observes as v; it is NOT the DB18 body velocity v_long.
Both reference _lane_frame variants are preserved: state.py (normalize iff norm > 0, right NOT normalized) vs continuous_state.py (normalize iff norm > 1e-12 else zeros, right normalized). Do not merge them.
crossing_available follows DuckController.crossing_available(i): (limit <= 0 || crossings_started[i] < limit) && crossing_armed[i].
Known deviations (documented, not defects)
acosconditioning × libm sin/cos (same class as the FJ2atan2deviation).get_lane_pos2.angle_rad = acos(dot(get_dir_vec(angle), tangent)); glibc (fixture) and OpenLibm (Windows Julia) sin/cos differ by ≤ 1 ULP, andacosamplifies this by1/|sin(angle)|near alignment. Measured worst case: 2.1e-14 rad atangle = 5.4e-3(11/300 rollout rows above 2 ULP;distbit-exact on every such row). The test comparator allows ≤ 2 ULP or the conditioning-scaled bound4·eps/max(|sin(angle)|, 1e-6).
FJ3 closing summary
Full suite after FJ3.8: 80 583 assertions, 0 failures, 0 errors (Pkg.test exit 0; verified from the complete log, not a truncated pipe).
FJ3.1 Map / collision / spawn geometry PASSED
FJ3.2 DB18 delayed dynamics PASSED
FJ3.3 Delay-buffer semantics PASSED
FJ3.4 Duckie walk dynamics PASSED
FJ3.5 DuckController.before_step PASSED
FJ3.6 State observers (f_tab / f_cont) PASSED
FJ3.7 Full one-decision transition chain PASSED
FJ3.8 RNG compatibility (A / B / C exact) PASSEDNext gate: FJ4 — the POMDPs.jl interface. gen becomes a thin adapter over simulate_decision ((sp = r.sp, r = r.reward.total)), plus initialstate built on the FJ3.1 spawn sampler with NumpyPCG64 available for reference-identical resets, and the observation-space projections. Carry the stop-sign scenario observation below into that gate.
Scenario observations
- ~~The injected stop sign may be geometrically unreachable as a candidate.~~ SUPERSEDED by the FJ5.4 probe — this note was wrong. The original observation was that with the reference qlearning/sarsa config (`stopspawnpos: [1.20, 2.10]
, rotate 180 → world(0.702, 1.8135))dstopstayedNoneacross a 300-decision rollout. That was an artifact of ONE trajectory (its spawn, policy and length), not a property of the configuration: probing the LIVE reference runtime for 400 decisions returns a candidate on 19 of them (dstopfrom 0.336 m down to 0.181 m as the ego approaches, minimum lateral offset 0.0023 m), verdictAREACHABLE. Seedocs/src/validation/FJ5STATUS.md§FJ5.4 and the raw data indocs/src/validation/fj54stop_probe.json`. The baseline config is fine and must not be changed.
FJ3.7 design
Canonical native transition: simulate_decision(model, s, action, rng) → TransitionResult(sp, raw_state, continuous_state, reward, events, terminated, truncated, reason, wheel_commands). POMDPs.gen (FJ5) will be a thin adapter returning (sp = r.sp, r = r.reward.total) — diagnostics, parity, visualization and tree logging all need the rich result.
RNG design (per the FJ3.8 decision): stochasticity is EXTERNAL — simulate_decision(m, s, a, rng) draws the p_cross trigger from the caller's rng via before_step(world, cfg, rng), and draws it ONLY on a fully eligible duck (RNG-A call-semantics parity). DuckieWorldState.controller_rng remains only for the legacy 2-arg before_step (stream-in-state) until FJ3.8 splits the reference-replay backend state; the transition chain never touches it. RNG-B (same state/action/seed → identical result) is a pinned test.
Acceptance evidence (all in test_fj3_transition.jl, fixture generated by the real wrapper step() functions — never a hand-reconstructed chain):
ONE-DECISION CHAIN PARITY 59 870 assertions over 1 250 decisions
Discrete actions (7 macro + mixed + variants) PASS
Continuous actions (10 key points + follow) PASS
world transition / raw state / StopTracker PASS
events / reward breakdown (10 comp) / reason PASS
terminated vs truncated PASS
Branch purity + aliasing (ego/history/q0/v0/
ducks/corners/norm/counters, sp1 vs sp2) PASS (35)
RNG-B determinism PASS
Frame-skip = 6 per decision PASS
Delay window serves the next decision's 1st tick PASS
Parent-state mutation NONEEvery TerminationReason and every EventFlags bit is reached: baseline scenarios give inprogress/offroad/timeout; explicit config-variant scenarios (recorded in scenario meta, baseline configs untouched) give goal (goaltile set), fullstop/passedstop/stopviolation (sign moved onto the route ring), othercollision (sign on the lane), and duck_collision (duckie spawned on the lane heading against travel).
FJ3.8 design and result — RNG-C EXACT
Compatibility layer only: the canonical model is untouched. NumpyMT19937 <: Random.AbstractRNG, so RNG-C is simply a choice of the rng argument — simulate_decision(m, s, a, NumpyMT19937(seed)) replays the reference controller stream, while any other AbstractRNG keeps the native semantics.
| Level | Meaning | Result |
|---|---|---|
| RNG-A | call semantics: a draw happens exactly when the reference draws | PASS (before_step draws only on a fully eligible duck; pinned by the decision-by-decision draw counts in part C) |
| RNG-B | same model + state + action + seed → same Julia result | PASS (pinned in test_fj3_transition.jl) |
| RNG-C | exact NumPy stream identity | PASS (exact) — both streams bit-for-bit |
Evidence (test_fj3_rng.jl, fixture from numpy 1.20.0 in ddm-ref):
A MT19937 (RandomState) 891 assertions, 9 seeds incl. 2^31-1 / 2^32-1
624-word state after seeding bit-exact
tempered uint32 stream bit-exact
random_sample() doubles bit-exact
B SeedSequence + PCG64 2 241 assertions, same 9 seeds
generate_state(8, uint64) bit-exact
PCG64 raw uint64 bit-exact
Generator.random() bit-exact
Generator.uniform(a, b) bit-exact
Generator.integers(0, n), 6 ranges bit-exact
Generator.standard_normal() bit-exact (ziggurat)
C p_cross = 0.5 trigger chain 1 402 assertions, 200 decisions
draw values / positions / count exact
activation pattern + trajectory exact
D Simulator.reset call log 180 assertions, 89 calls
uniform / integers / normal bit-exact per callImplementation notes (each was a real discrepancy found by the fixtures):
SeedSequence.mixisMIX_MULT_L*x − MIX_MULT_R*y(uint32 wrapping SUBTRACTION), not a xor-combine. Verified against numpy v1.20.0bit_generator.pyx(hashmix/mix).Generator.integerswith a range that fits in 32 bits usesbuffered_bounded_lemire_uint32overpcg64_next32, not the 64-bit Lemire:next32returns the LOW half of a fresh uint64 and buffers the HIGH half in the bit-generator state, so the buffer survives across calls (it is part ofPCG64.state:has_uint32/uinteger). The double and uint64 paths bypass the buffer without clearing it.NumpyPCG64carries both fields for this reason.standard_normalis the numpy ziggurat; the tables were extracted verbatim fromziggurat_constants.hintosrc/rng/ziggurat_constants.jl. The ~99.3% fast path is integer × table (bit-exact); the wedge/tail paths use libmlog/expand inherit the documented ≤1-ULP cross-libm caveat.- uint64 fixture values are stored as decimal strings — JSON numbers above 2^53 lose bits through the Float64 path in the Julia JSON reader.
Part D replays each logged Generator call from the reference bit-generator state recorded immediately before it (state/inc + buffered half-word). This isolates per-method semantics and does not require the log to capture every draw consumer inside Simulator.reset — a few unlogged Generator methods also advance the stream during map/object init. Pure stream identity from the seed is pinned independently by part B.
Real defects found and fixed during FJ3.7 (previously masked)
The original FJ3.4/3.5 "green" state was partially unverified: the fixture's missing visible key made duck_close throw before its corners/norm comparisons ever ran, and runtests.jl aborts at the first failing testset, so FJ3.5 and everything after never executed. Chain-level testing then exposed:
generate_normeigen convention (FJ3.4 state-parity bug). The analytic implementation ordered the covariance eigenvectors by eigenvalue and chose its own signs — equivalent for SAT booleans but NOT for the recordedobj_normstate (np.linalg.eigdoes not sort). Replaced by a directLAPACK.geev!(dgeev) call, the same driver NumPy uses; the duckie fixtures now compare norm bit-level (≤ 2 ULP) through full walk cycles._valid_posedouble_actual_centeroffset (reference quirk).Simulator._valid_posere-assignspos = _actual_center(pos, angle)and then callsget_agent_corners(pos, angle)on the already-centered pos, so its collision box is shifted ~0.024 m behind the true agent box. The wrapper's_duck_collisionuses the true box, so at contact the two disagree for a decision or two and termination waits for the shifted box (or drivability) to fail. The Julia port now reproduces the double offset verbatim (see the comment in_valid_pose); without it, collisions terminated exactly one decision early.before_steptuple/Vector mismatch (closest_curve_pointrequires a Vector) — latentMethodErroron every eligible trigger evaluation.walk_distancestamping.duck_initinfj3_duck.jsonis recorded BEFOREDuckController.__init__stampscfg.walk_distance(0.90) onto the duck, so it carries the map default (0.585). The FJ3.5 rollout world must use the stamped value (the Part-A walk scenarios use the recorded default). Also the rollout ego starts at the recorded reset pose (rollout[0].pos_pre), not at the duck position.
Fix log (2026-08-19)
test_fj3_duck.jlleft failing by the previous session: (a) the fixture lacked thevisiblekey thatduck_closecompares — added"visible": bool(duck.visible)togen_fj3_duck_fixtures.pyand regenerated; (b) the tick-chain sanity assertion lacked the "both ticks active" case (a continuously-walking duck moves whileactivestays true) — the invariant is nowcenter may only change on a tick where the duck walked.- Full-suite verification is now logged to a file and checked for the final summary line — an earlier
| head -50pipeline truncated the output and masked late-running testset failures (the pipeline exit code ishead's).