FJ3 — Native Dynamics, Observers, Transition, RNG

Last update: 2026-08-19. Status: FJ3.1–FJ3.8 done — FJ3 COMPLETE.

RNG-C claim, in the wording to use in the README/paper:

Bit-exact compatibility for the tested NumPy RNG streams and exercised generator paths; rare libm-dependent normal-sampling branches retain the documented cross-platform ≤1-ULP caveat.

(The ziggurat fast path — ~99.3% of standard_normal draws — is integer × table and therefore bit-exact; only the wedge/tail rejection branches call log/exp. Nothing in the MDP transition chain draws normals at all: the controller stream uses random_sample only.)

FJ3 ports the latent dynamics and the env-dependent state extraction that FJ2 could not pin as pure functions. Every substep is pinned against fixtures generated by the real reference stack in the ddm-ref conda env (Python 3.9, numpy 1.20.0, gym 0.23.1, duckietown-gym-daffy 6.1.34, duckietown-world-daffy 6.4.3) plus the duckduck/src modules.

Substep table

SubstepScopeFixture / testStatus
FJ3.1Map reconstruction (small_loop tiles, Bézier lanes, closest_curve_point, get_lane_pos2, collision geometry, spawn sampling)fj3_map.json / test_fj3_map.jldone
FJ3.2DB18 nominal motor model + 0.15 s delayed dynamics, SE(2) chain, weird_from_cartesianfj3_ego.json / test_fj3_ego.jldone
FJ3.3Delay-buffer window semantics (trim, u0 fallback, cross-decision memory)fj3_ego.json / test_fj3_delay.jldone
FJ3.4DuckieObj.step walk / finish_walk dynamics (600-tick standalone scenarios)fj3_duck.json / test_fj3_duck.jldone
FJ3.5DuckController.before_step trigger semantics + full 137-decision rollout chain parity (ego ticks + duck ticks + trigger)fj3_duck.json / test_fj3_duck.jldone
FJ3.6State-extraction observers: get_raw_state, next_stop_candidate, classify_duck, tile_ahead, both _lane_frame variants, duck_relative_state, signed_curvature_ahead, get_continuous_state, encodefj3_obs.json (334 rows) / test_fj3_obs.jl (18 994 assertions)done
FJ3.7One-decision transition chain (simulate_decision → TransitionResult): before_step → wheels → frame_skip ticks → sim-done → extraction → StopTracker → collision/reason → terminated/truncated → reward; discrete AND continuousfj37_transition.json (24 scenarios, 1250 decisions, generated by the REAL DuckieMDPEnv.step/ContinuousDuckieMDPEnv.step) / test_fj3_transition.jl (59 913 assertions)done
FJ3.8Exact NumPy RNG stream identity: NumpyMT19937 (= np.random.RandomState) and NumpySeedSequence/NumpyPCG64 (= gym.utils.seeding.np_random) + the Generator methods the reset path callsfj38_rng.json / test_fj3_rng.jl (4 714 assertions)done — RNG-C exact

FJ3.6 design

Fixture rows are latent-in → outputs-out: each row records the full latent input (ego pose/speed, duck center/heading/vel/active/visible, controller counters, sigma/fallback memory) and every observer output. The Julia test rebuilds a DuckieWorldState from the row, so extraction parity is pinned without accumulating dynamics drift (chain-level integration is FJ3.7). 300 rows from a driven 3-lap rollout + 34 synthetic rows covering: every next_stop_candidate filter branch (distance/behind/lateral/ orientation), all 5 DuckThreat classes, all 3 ego-relative tile classes, the detection gate, controller-counter crossing_available, the NotInLane fallback, and the off-map STRAIGHT fallback.

New state field: DuckieEgoState.speed — the simulator's self.speed = norm(cur_pos - prev_pos) / delta_time (displacement speed of the last tick). This is what get_raw_state observes as v; it is NOT the DB18 body velocity v_long.

Both reference _lane_frame variants are preserved: state.py (normalize iff norm > 0, right NOT normalized) vs continuous_state.py (normalize iff norm > 1e-12 else zeros, right normalized). Do not merge them.

crossing_available follows DuckController.crossing_available(i): (limit <= 0 || crossings_started[i] < limit) && crossing_armed[i].

Known deviations (documented, not defects)

  1. acos conditioning × libm sin/cos (same class as the FJ2 atan2 deviation). get_lane_pos2.angle_rad = acos(dot(get_dir_vec(angle), tangent)); glibc (fixture) and OpenLibm (Windows Julia) sin/cos differ by ≤ 1 ULP, and acos amplifies this by 1/|sin(angle)| near alignment. Measured worst case: 2.1e-14 rad at angle = 5.4e-3 (11/300 rollout rows above 2 ULP; dist bit-exact on every such row). The test comparator allows ≤ 2 ULP or the conditioning-scaled bound 4·eps/max(|sin(angle)|, 1e-6).

FJ3 closing summary

Full suite after FJ3.8: 80 583 assertions, 0 failures, 0 errors (Pkg.test exit 0; verified from the complete log, not a truncated pipe).

FJ3.1 Map / collision / spawn geometry     PASSED
FJ3.2 DB18 delayed dynamics                PASSED
FJ3.3 Delay-buffer semantics               PASSED
FJ3.4 Duckie walk dynamics                 PASSED
FJ3.5 DuckController.before_step           PASSED
FJ3.6 State observers (f_tab / f_cont)     PASSED
FJ3.7 Full one-decision transition chain   PASSED
FJ3.8 RNG compatibility (A / B / C exact)  PASSED

Next gate: FJ4 — the POMDPs.jl interface. gen becomes a thin adapter over simulate_decision ((sp = r.sp, r = r.reward.total)), plus initialstate built on the FJ3.1 spawn sampler with NumpyPCG64 available for reference-identical resets, and the observation-space projections. Carry the stop-sign scenario observation below into that gate.

Scenario observations

  1. ~~The injected stop sign may be geometrically unreachable as a candidate.~~ SUPERSEDED by the FJ5.4 probe — this note was wrong. The original observation was that with the reference qlearning/sarsa config (`stopspawnpos: [1.20, 2.10], rotate 180 → world(0.702, 1.8135))dstopstayedNoneacross a 300-decision rollout. That was an artifact of ONE trajectory (its spawn, policy and length), not a property of the configuration: probing the LIVE reference runtime for 400 decisions returns a candidate on 19 of them (dstopfrom 0.336 m down to 0.181 m as the ego approaches, minimum lateral offset 0.0023 m), verdictAREACHABLE. Seedocs/src/validation/FJ5STATUS.md§FJ5.4 and the raw data indocs/src/validation/fj54stop_probe.json`. The baseline config is fine and must not be changed.

FJ3.7 design

Canonical native transition: simulate_decision(model, s, action, rng) → TransitionResult(sp, raw_state, continuous_state, reward, events, terminated, truncated, reason, wheel_commands). POMDPs.gen (FJ5) will be a thin adapter returning (sp = r.sp, r = r.reward.total) — diagnostics, parity, visualization and tree logging all need the rich result.

RNG design (per the FJ3.8 decision): stochasticity is EXTERNAL — simulate_decision(m, s, a, rng) draws the p_cross trigger from the caller's rng via before_step(world, cfg, rng), and draws it ONLY on a fully eligible duck (RNG-A call-semantics parity). DuckieWorldState.controller_rng remains only for the legacy 2-arg before_step (stream-in-state) until FJ3.8 splits the reference-replay backend state; the transition chain never touches it. RNG-B (same state/action/seed → identical result) is a pinned test.

Acceptance evidence (all in test_fj3_transition.jl, fixture generated by the real wrapper step() functions — never a hand-reconstructed chain):

ONE-DECISION CHAIN PARITY          59 870 assertions over 1 250 decisions
  Discrete actions (7 macro + mixed + variants)   PASS
  Continuous actions (10 key points + follow)     PASS
  world transition / raw state / StopTracker      PASS
  events / reward breakdown (10 comp) / reason    PASS
  terminated vs truncated                         PASS
Branch purity + aliasing (ego/history/q0/v0/
  ducks/corners/norm/counters, sp1 vs sp2)        PASS (35)
RNG-B determinism                                 PASS
Frame-skip = 6 per decision                       PASS
Delay window serves the next decision's 1st tick  PASS
Parent-state mutation                             NONE

Every TerminationReason and every EventFlags bit is reached: baseline scenarios give inprogress/offroad/timeout; explicit config-variant scenarios (recorded in scenario meta, baseline configs untouched) give goal (goaltile set), fullstop/passedstop/stopviolation (sign moved onto the route ring), othercollision (sign on the lane), and duck_collision (duckie spawned on the lane heading against travel).

FJ3.8 design and result — RNG-C EXACT

Compatibility layer only: the canonical model is untouched. NumpyMT19937 <: Random.AbstractRNG, so RNG-C is simply a choice of the rng argument — simulate_decision(m, s, a, NumpyMT19937(seed)) replays the reference controller stream, while any other AbstractRNG keeps the native semantics.

LevelMeaningResult
RNG-Acall semantics: a draw happens exactly when the reference drawsPASS (before_step draws only on a fully eligible duck; pinned by the decision-by-decision draw counts in part C)
RNG-Bsame model + state + action + seed → same Julia resultPASS (pinned in test_fj3_transition.jl)
RNG-Cexact NumPy stream identityPASS (exact) — both streams bit-for-bit

Evidence (test_fj3_rng.jl, fixture from numpy 1.20.0 in ddm-ref):

A  MT19937 (RandomState)      891 assertions, 9 seeds incl. 2^31-1 / 2^32-1
     624-word state after seeding        bit-exact
     tempered uint32 stream              bit-exact
     random_sample() doubles             bit-exact
B  SeedSequence + PCG64      2 241 assertions, same 9 seeds
     generate_state(8, uint64)           bit-exact
     PCG64 raw uint64                    bit-exact
     Generator.random()                  bit-exact
     Generator.uniform(a, b)             bit-exact
     Generator.integers(0, n), 6 ranges  bit-exact
     Generator.standard_normal()         bit-exact (ziggurat)
C  p_cross = 0.5 trigger chain 1 402 assertions, 200 decisions
     draw values / positions / count     exact
     activation pattern + trajectory     exact
D  Simulator.reset call log      180 assertions, 89 calls
     uniform / integers / normal         bit-exact per call

Implementation notes (each was a real discrepancy found by the fixtures):

  1. SeedSequence.mix is MIX_MULT_L*x − MIX_MULT_R*y (uint32 wrapping SUBTRACTION), not a xor-combine. Verified against numpy v1.20.0 bit_generator.pyx (hashmix/mix).
  2. Generator.integers with a range that fits in 32 bits uses buffered_bounded_lemire_uint32 over pcg64_next32, not the 64-bit Lemire: next32 returns the LOW half of a fresh uint64 and buffers the HIGH half in the bit-generator state, so the buffer survives across calls (it is part of PCG64.state: has_uint32/uinteger). The double and uint64 paths bypass the buffer without clearing it. NumpyPCG64 carries both fields for this reason.
  3. standard_normal is the numpy ziggurat; the tables were extracted verbatim from ziggurat_constants.h into src/rng/ziggurat_constants.jl. The ~99.3% fast path is integer × table (bit-exact); the wedge/tail paths use libm log/exp and inherit the documented ≤1-ULP cross-libm caveat.
  4. uint64 fixture values are stored as decimal strings — JSON numbers above 2^53 lose bits through the Float64 path in the Julia JSON reader.

Part D replays each logged Generator call from the reference bit-generator state recorded immediately before it (state/inc + buffered half-word). This isolates per-method semantics and does not require the log to capture every draw consumer inside Simulator.reset — a few unlogged Generator methods also advance the stream during map/object init. Pure stream identity from the seed is pinned independently by part B.

Real defects found and fixed during FJ3.7 (previously masked)

The original FJ3.4/3.5 "green" state was partially unverified: the fixture's missing visible key made duck_close throw before its corners/norm comparisons ever ran, and runtests.jl aborts at the first failing testset, so FJ3.5 and everything after never executed. Chain-level testing then exposed:

  1. generate_norm eigen convention (FJ3.4 state-parity bug). The analytic implementation ordered the covariance eigenvectors by eigenvalue and chose its own signs — equivalent for SAT booleans but NOT for the recorded obj_norm state (np.linalg.eig does not sort). Replaced by a direct LAPACK.geev! (dgeev) call, the same driver NumPy uses; the duckie fixtures now compare norm bit-level (≤ 2 ULP) through full walk cycles.
  2. _valid_pose double _actual_center offset (reference quirk). Simulator._valid_pose re-assigns pos = _actual_center(pos, angle) and then calls get_agent_corners(pos, angle) on the already-centered pos, so its collision box is shifted ~0.024 m behind the true agent box. The wrapper's _duck_collision uses the true box, so at contact the two disagree for a decision or two and termination waits for the shifted box (or drivability) to fail. The Julia port now reproduces the double offset verbatim (see the comment in _valid_pose); without it, collisions terminated exactly one decision early.
  3. before_step tuple/Vector mismatch (closest_curve_point requires a Vector) — latent MethodError on every eligible trigger evaluation.
  4. walk_distance stamping. duck_init in fj3_duck.json is recorded BEFORE DuckController.__init__ stamps cfg.walk_distance (0.90) onto the duck, so it carries the map default (0.585). The FJ3.5 rollout world must use the stamped value (the Part-A walk scenarios use the recorded default). Also the rollout ego starts at the recorded reset pose (rollout[0].pos_pre), not at the duck position.

Fix log (2026-08-19)

  • test_fj3_duck.jl left failing by the previous session: (a) the fixture lacked the visible key that duck_close compares — added "visible": bool(duck.visible) to gen_fj3_duck_fixtures.py and regenerated; (b) the tick-chain sanity assertion lacked the "both ticks active" case (a continuously-walking duck moves while active stays true) — the invariant is now center may only change on a tick where the duck walked.
  • Full-suite verification is now logged to a file and checked for the final summary line — an earlier | head -50 pipeline truncated the output and masked late-running testset failures (the pipeline exit code is head's).