FJ9.6 — Diagnostic Time Series

Date: 2026-08-23.

The renderer reads the frozen FJ8.4c decision log and nothing else. No environment, no policy, no planner, and no solver runs anywhere in this gate.

STATUS

PASSED. FJ9.6a–9.6d closed. The decision-artifact contract is executable, the DiagnosticSeries representation carries unit, category, kind and semantics from the core, five sections render for four episodes across three solver families, and the aggregate reproduces the FJ8.4c compute-collapse table from an independently written binning.

One documentation defect in FJ8 was found and corrected. See NEW DEFECTS.

ENVIRONMENT

WSL Ubuntu-Baru, juliaup override 1.11.3
Pkg.test with ddm-ref + ddm-torch + MCTS.jl : 173 testsets, 148 471 assertions
render check (fresh process, CairoMakie only): RENDER_EXIT=0
                                               PYTHON_MODULES=none
                                               MCTS_LOADED=false

FJ9.6a — THE DECISION-ARTIFACT CONTRACT

The audit is code, not a claim: decision_log_audit probes the artefact against DECISION_QUANTITY_CONTRACT and reports one line per wanted quantity. If a future artefact drops a column, the line changes from LOGGED to ABSENT on its own.

source          artifacts/fj8/enriched/decisions.csv
fingerprint     952e2a164018018a
rows            16 522 decisions over 120 episodes (6 solvers × 20 seeds)
LOGGED          42 columns
DERIVED         2  (exact identities, labelled as derived)
ABSENT          4

Every quantity FJ9.6a asked for is present. The full table is artifacts/fj9/decision_contract.md, regenerated by the render check.

What is ABSENT, and stays absent: the observation as the policy saw it, the belief state, the per-decision search tree, and a wall-clock timestamp. None is back-computed. FJ9.5 persists search trees only for the two snapshots that were explicitly captured, so a per-decision tree panel is not available and is not faked.

The one column with missing values is d_stop — 12 814 of 16 522 rows, 77.6 %. It is missing exactly when no stop sign is a candidate. This is the only missingness in the artefact, and the audit reports it as a count rather than as prose.

Two derived quantities, both exact identities:

DerivedIdentityWhy it is not an inference
cumulative_returnrunning sum of reward_totalFJ8.4c showed the final value equals the FJ8.4b episode return exactly
horizon_reachedterminated == false on the final rowmeasured; see below

truncated carries no signal. It is false on all 16 522 rows. The evaluator's horizon stop is not recorded as truncation, so the outcome is read from terminated instead, and the loader validates the consequence: of 120 episodes, 99 end with terminated=false, reason=in_progress at exactly 150 decisions, and 21 end with terminated=true before it (15 offroad, 6 other_collision). A renderer that had trusted truncated would have labelled every one of the 99 as a normal completion.

FJ9.6b — THE CORE REPRESENTATION

decisions.csv -> load_decision_log -> DecisionLog -> episode_diagnostics
                                                  -> EpisodeDiagnostics
                                                     -> render_diagnostics

load_decision_log refuses what it cannot interpret, before any figure exists: a missing required column (named in the error), a ragged row, a gap in an episode's decision indices, or a terminated flag before the final decision. An episode with a hole in it is a corrupt record, not a short episode.

Each DiagnosticSeries carries name, values, unit, category, kind and semantics. A backend never decides whether a column is metres, radians, a flag or a cumulative quantity — that decision is the model's, and FJ9.2 established that it belongs in the core.

missing survives end to end. The three cases stay three cases: missing, exactly 0.0, and a real value. A test asserts they are not interchangeable.

Two x-axis modes exist and a figure must state which it uses: ABSOLUTE_DECISION (index 1..T) and NORMALIZED_PROGRESS ((k−1)/T). The fingerprint includes the mode, so the same episode in the two modes is two different figures. Normalised progress is not time, and the axis label says so.

Event markers come from episode.events, which holds the decision indices the log recorded. Nothing is inferred from a threshold — FJ9.1 showed how convincing an inferred marker looks while corresponding to nothing.

FJ9.6c — SINGLE-EPISODE RENDERING

Six sections: Navigation, Motion and command, Stop subsystem, Duck subsystem, Reward, Computational cost.

Deviation from the requested five panels, stated deliberately: within a section there is one axis per unit. The first draft followed the five-panel grouping literally and it produced exactly the failure the grouping was meant to prevent. On the Navigation panel, kappa reaches 2.15 (1/m) while d never leaves ±0.15 m, so the lateral offset rendered as a flat line. On the compute panel, planning_time at 0.1 s beside model_calls at 1 400 read as zero for the whole episode. The section is the reader's grouping; the axis is the physical one, and merging them costs the scale of everything but the largest series. The sections are unchanged; they now contain 11 axes.

Flags (sigma_stop, duck_present, duck_active) are shaded across every axis of their section rather than drawn as lines, because a flag is a state the episode was in, not a value with a scale.

Rendered from the frozen log, in a fresh process, with no Python and no MCTS:

EpisodeWhy
mcts@1k seed 1001a planner that survives to the horizon
dpw@1k seed 1001the compute collapse, decision by decision
td3 seed 1001a learned continuous policy
q_learning seed 1001model_calls = 0 as a measured zero; the only one of the four whose flags actually fire

PNG and SVG both produced; diagnostics_dpw1k_1001_progress.png is the same episode on the normalised axis, and its fingerprint differs from the absolute one.

FJ9.6d — AGGREGATE DIAGNOSTICS

progress_bins aggregates by bin membership only. There is no interpolation: stretching a 42-decision DPW episode onto a 150-point grid would manufacture values that never occurred. Every bin reports its own n, and the figure prints it.

Mean generative calls per decision, dpw@1k, by normalised progress:

bin          0-20%   20-40%   40-60%   60-80%   80-100%
mean          1019      835      940      602       195
n              350      342      343      342       336

These are the FJ8.4c numbers, reproduced by a binning written independently of tools/enrich_decision_log.jl. Two implementations agreeing is a check on both; one agreeing with itself is not. The test pins the five values.

q_learning bins to exactly 0 in every bin — measured, not missing.

A column with missing values reports how many rows it excluded rather than substituting zeros: d_stop over dpw@1k excludes 1 224 of its 1 713 rows and prints that on the figure.

FILES

FilePurpose
src/visualization/diagnostics.jlthe whole core: loader, audit, series, bins
src/visualization/scene.jlrender_diagnostics, render_diagnostics_aggregate declarations
ext/DuckietownMakieExt.jlthe drawing half
src/interfaces/pomdp_readiness.jlrender_diagnostics added to the extension points (now 8)
test/test_fj96_diagnostics.jl452 assertions across 11 testsets
tools/render_check.jlfresh-process render, PNG + SVG
artifacts/fj9/decision_contract.mdthe FJ9.6a table
artifacts/fj9/diagnostics_*.png/.svgeight figures

NEGATIVE CONTROL

Edit one model_calls cell in a copy of the log. Required outcome: the computational series and the fingerprints move, and nothing else does.

source_fingerprint          changed
diagnostics_fingerprint     changed
model_calls series          changed (999.0 at the edited decision)
d, phi, v_cmd, reward_total,
cumulative_return, d_stop,
planning_time               byte-identical
events, outcome             unchanged

KNOWN DEVIATIONS

  1. Six sections, eleven axes, rather than five single-axis panels. Deliberate; the reason is measured and given above.
  2. No per-decision search-tree panel. The evaluation never stored trees per decision. Reported ABSENT.
  3. duck_class and duck_active_state are logged but not plotted. They are categorical; the contract records them as available for a later gate.

NEW DEFECTS

One, in FJ8's prose, found by this gate's figures and corrected.

docs/src/validation/FJ8_STATUS.md stated that TD3 "never reaches a stop sign". The TD3 time series shows the opposite: full_stop fires at decision 28, sigma_stop is true for 123 of the 150 decisions that follow, d_stop falls to 0, and v sits at 0.005 m/s for the rest of the episode. TD3 reaches the sign, stops, and never proceeds — in 20 of 20 evaluation seeds (passed_stops = 0). Its −242 return is −234 stagnation penalty.

The FJ8.4b artefact already contained the evidence — stop_zone_decisions = 2289 for TD3, about 114 per episode — so only the sentence was wrong, and the n/a compliance verdict it was attached to was right for a different reason than the one given: TD3 has no compliance rate because it completes no encounter, not because it never meets one. The paragraph is corrected in place with a dated note, no number is changed, and the claim is now pinned by a test.

Comparison across solvers, from the same log:

episodes that full_stop but never passed_stop
  td3          20/20
  dpw@1k        4/20
  mcts@1k       1/20
  sac           0/20
  q_learning    0/20
  sarsa         0/20

Two lesser bugs of my own, both caught by looking at the rendered output rather than at the exit code: axes holding stop_hold_progress and the reward components were labelled "flag" because their unit string is empty, and fixed row heights made the layout exceed the canvas so Makie silently clipped the title and the provenance footer off the top and bottom. Both fixed. A render check that only asserts filesize > 5_000 would have passed all three.

SCIENTIFIC INTERPRETATION

The gate was justified by what an aggregate cannot say, and it produced two examples on its first run.

DPW's compute collapse was already known from FJ8.4c as five binned means. The per-decision series shows the mechanism: consumption tracks trajectory quality within the episode, rising to ~1 400 calls around the stop encounter at decision 52 and decaying to 24 by the offroad termination at 116. The budget is endogenous, and the shape of the decay is the planner running out of tree to expand as its state deteriorates.

TD3 is the sharper case. Every aggregate metric describing it was correct, and the picture they gave was still wrong: 0 offroad, 0 collisions, 0 violations, a full stop in every episode, and compliance reported honestly as n/a. Read as a row in a table, that is a cautious policy. Read as a time series, it is a policy that stops at the first sign and stalls there for the remaining 120 decisions. The n/a was the denominator rule working — but it was the only signal in the table that something was wrong, and it is easy to read as "insufficient data" rather than "never finished".

Nothing new was measured to find this. The evidence was in the FJ8.4b artefact in June and in the FJ8.4c log last week. What changed is that the record is now displayed at the resolution at which the behaviour exists.

NEXT GATE

FJ9.7 — animation.