FJ9.6 — Diagnostic Time Series
Date: 2026-08-23.
The renderer reads the frozen FJ8.4c decision log and nothing else. No environment, no policy, no planner, and no solver runs anywhere in this gate.
STATUS
PASSED. FJ9.6a–9.6d closed. The decision-artifact contract is executable, the DiagnosticSeries representation carries unit, category, kind and semantics from the core, five sections render for four episodes across three solver families, and the aggregate reproduces the FJ8.4c compute-collapse table from an independently written binning.
One documentation defect in FJ8 was found and corrected. See NEW DEFECTS.
ENVIRONMENT
WSL Ubuntu-Baru, juliaup override 1.11.3
Pkg.test with ddm-ref + ddm-torch + MCTS.jl : 173 testsets, 148 471 assertions
render check (fresh process, CairoMakie only): RENDER_EXIT=0
PYTHON_MODULES=none
MCTS_LOADED=falseFJ9.6a — THE DECISION-ARTIFACT CONTRACT
The audit is code, not a claim: decision_log_audit probes the artefact against DECISION_QUANTITY_CONTRACT and reports one line per wanted quantity. If a future artefact drops a column, the line changes from LOGGED to ABSENT on its own.
source artifacts/fj8/enriched/decisions.csv
fingerprint 952e2a164018018a
rows 16 522 decisions over 120 episodes (6 solvers × 20 seeds)
LOGGED 42 columns
DERIVED 2 (exact identities, labelled as derived)
ABSENT 4Every quantity FJ9.6a asked for is present. The full table is artifacts/fj9/decision_contract.md, regenerated by the render check.
What is ABSENT, and stays absent: the observation as the policy saw it, the belief state, the per-decision search tree, and a wall-clock timestamp. None is back-computed. FJ9.5 persists search trees only for the two snapshots that were explicitly captured, so a per-decision tree panel is not available and is not faked.
The one column with missing values is d_stop — 12 814 of 16 522 rows, 77.6 %. It is missing exactly when no stop sign is a candidate. This is the only missingness in the artefact, and the audit reports it as a count rather than as prose.
Two derived quantities, both exact identities:
| Derived | Identity | Why it is not an inference |
|---|---|---|
cumulative_return | running sum of reward_total | FJ8.4c showed the final value equals the FJ8.4b episode return exactly |
horizon_reached | terminated == false on the final row | measured; see below |
truncated carries no signal. It is false on all 16 522 rows. The evaluator's horizon stop is not recorded as truncation, so the outcome is read from terminated instead, and the loader validates the consequence: of 120 episodes, 99 end with terminated=false, reason=in_progress at exactly 150 decisions, and 21 end with terminated=true before it (15 offroad, 6 other_collision). A renderer that had trusted truncated would have labelled every one of the 99 as a normal completion.
FJ9.6b — THE CORE REPRESENTATION
decisions.csv -> load_decision_log -> DecisionLog -> episode_diagnostics
-> EpisodeDiagnostics
-> render_diagnosticsload_decision_log refuses what it cannot interpret, before any figure exists: a missing required column (named in the error), a ragged row, a gap in an episode's decision indices, or a terminated flag before the final decision. An episode with a hole in it is a corrupt record, not a short episode.
Each DiagnosticSeries carries name, values, unit, category, kind and semantics. A backend never decides whether a column is metres, radians, a flag or a cumulative quantity — that decision is the model's, and FJ9.2 established that it belongs in the core.
missing survives end to end. The three cases stay three cases: missing, exactly 0.0, and a real value. A test asserts they are not interchangeable.
Two x-axis modes exist and a figure must state which it uses: ABSOLUTE_DECISION (index 1..T) and NORMALIZED_PROGRESS ((k−1)/T). The fingerprint includes the mode, so the same episode in the two modes is two different figures. Normalised progress is not time, and the axis label says so.
Event markers come from episode.events, which holds the decision indices the log recorded. Nothing is inferred from a threshold — FJ9.1 showed how convincing an inferred marker looks while corresponding to nothing.
FJ9.6c — SINGLE-EPISODE RENDERING
Six sections: Navigation, Motion and command, Stop subsystem, Duck subsystem, Reward, Computational cost.
Deviation from the requested five panels, stated deliberately: within a section there is one axis per unit. The first draft followed the five-panel grouping literally and it produced exactly the failure the grouping was meant to prevent. On the Navigation panel, kappa reaches 2.15 (1/m) while d never leaves ±0.15 m, so the lateral offset rendered as a flat line. On the compute panel, planning_time at 0.1 s beside model_calls at 1 400 read as zero for the whole episode. The section is the reader's grouping; the axis is the physical one, and merging them costs the scale of everything but the largest series. The sections are unchanged; they now contain 11 axes.
Flags (sigma_stop, duck_present, duck_active) are shaded across every axis of their section rather than drawn as lines, because a flag is a state the episode was in, not a value with a scale.
Rendered from the frozen log, in a fresh process, with no Python and no MCTS:
| Episode | Why |
|---|---|
mcts@1k seed 1001 | a planner that survives to the horizon |
dpw@1k seed 1001 | the compute collapse, decision by decision |
td3 seed 1001 | a learned continuous policy |
q_learning seed 1001 | model_calls = 0 as a measured zero; the only one of the four whose flags actually fire |
PNG and SVG both produced; diagnostics_dpw1k_1001_progress.png is the same episode on the normalised axis, and its fingerprint differs from the absolute one.
FJ9.6d — AGGREGATE DIAGNOSTICS
progress_bins aggregates by bin membership only. There is no interpolation: stretching a 42-decision DPW episode onto a 150-point grid would manufacture values that never occurred. Every bin reports its own n, and the figure prints it.
Mean generative calls per decision, dpw@1k, by normalised progress:
bin 0-20% 20-40% 40-60% 60-80% 80-100%
mean 1019 835 940 602 195
n 350 342 343 342 336These are the FJ8.4c numbers, reproduced by a binning written independently of tools/enrich_decision_log.jl. Two implementations agreeing is a check on both; one agreeing with itself is not. The test pins the five values.
q_learning bins to exactly 0 in every bin — measured, not missing.
A column with missing values reports how many rows it excluded rather than substituting zeros: d_stop over dpw@1k excludes 1 224 of its 1 713 rows and prints that on the figure.
FILES
| File | Purpose |
|---|---|
src/visualization/diagnostics.jl | the whole core: loader, audit, series, bins |
src/visualization/scene.jl | render_diagnostics, render_diagnostics_aggregate declarations |
ext/DuckietownMakieExt.jl | the drawing half |
src/interfaces/pomdp_readiness.jl | render_diagnostics added to the extension points (now 8) |
test/test_fj96_diagnostics.jl | 452 assertions across 11 testsets |
tools/render_check.jl | fresh-process render, PNG + SVG |
artifacts/fj9/decision_contract.md | the FJ9.6a table |
artifacts/fj9/diagnostics_*.png/.svg | eight figures |
NEGATIVE CONTROL
Edit one model_calls cell in a copy of the log. Required outcome: the computational series and the fingerprints move, and nothing else does.
source_fingerprint changed
diagnostics_fingerprint changed
model_calls series changed (999.0 at the edited decision)
d, phi, v_cmd, reward_total,
cumulative_return, d_stop,
planning_time byte-identical
events, outcome unchangedKNOWN DEVIATIONS
- Six sections, eleven axes, rather than five single-axis panels. Deliberate; the reason is measured and given above.
- No per-decision search-tree panel. The evaluation never stored trees per decision. Reported ABSENT.
duck_classandduck_active_stateare logged but not plotted. They are categorical; the contract records them as available for a later gate.
NEW DEFECTS
One, in FJ8's prose, found by this gate's figures and corrected.
docs/src/validation/FJ8_STATUS.md stated that TD3 "never reaches a stop sign". The TD3 time series shows the opposite: full_stop fires at decision 28, sigma_stop is true for 123 of the 150 decisions that follow, d_stop falls to 0, and v sits at 0.005 m/s for the rest of the episode. TD3 reaches the sign, stops, and never proceeds — in 20 of 20 evaluation seeds (passed_stops = 0). Its −242 return is −234 stagnation penalty.
The FJ8.4b artefact already contained the evidence — stop_zone_decisions = 2289 for TD3, about 114 per episode — so only the sentence was wrong, and the n/a compliance verdict it was attached to was right for a different reason than the one given: TD3 has no compliance rate because it completes no encounter, not because it never meets one. The paragraph is corrected in place with a dated note, no number is changed, and the claim is now pinned by a test.
Comparison across solvers, from the same log:
episodes that full_stop but never passed_stop
td3 20/20
dpw@1k 4/20
mcts@1k 1/20
sac 0/20
q_learning 0/20
sarsa 0/20Two lesser bugs of my own, both caught by looking at the rendered output rather than at the exit code: axes holding stop_hold_progress and the reward components were labelled "flag" because their unit string is empty, and fixed row heights made the layout exceed the canvas so Makie silently clipped the title and the provenance footer off the top and bottom. Both fixed. A render check that only asserts filesize > 5_000 would have passed all three.
SCIENTIFIC INTERPRETATION
The gate was justified by what an aggregate cannot say, and it produced two examples on its first run.
DPW's compute collapse was already known from FJ8.4c as five binned means. The per-decision series shows the mechanism: consumption tracks trajectory quality within the episode, rising to ~1 400 calls around the stop encounter at decision 52 and decaying to 24 by the offroad termination at 116. The budget is endogenous, and the shape of the decay is the planner running out of tree to expand as its state deteriorates.
TD3 is the sharper case. Every aggregate metric describing it was correct, and the picture they gave was still wrong: 0 offroad, 0 collisions, 0 violations, a full stop in every episode, and compliance reported honestly as n/a. Read as a row in a table, that is a cautious policy. Read as a time series, it is a policy that stops at the first sign and stalls there for the remaining 120 decisions. The n/a was the denominator rule working — but it was the only signal in the table that something was wrong, and it is easy to read as "insufficient data" rather than "never finished".
Nothing new was measured to find this. The evidence was in the FJ8.4b artefact in June and in the FJ8.4c log last week. What changed is that the record is now displayed at the resolution at which the behaviour exists.
NEXT GATE
FJ9.7 — animation.