Evaluation

Parity, rollouts, metrics, budget and comparison harnesses.

Duckietown.WorldMismatch — Type
WorldMismatch

One field where two world states differ, with both values rendered for the report. Floats are compared bitwise (===), so 0.0 and -0.0 count as different and NaN equals NaN — a planner-purity check must not be fooled by numeric coincidence.

source
Duckietown.benchmark_gen — Method
benchmark_gen(mdp, s0, a; calls, mode, seed) -> GenBenchmark

Cost of POMDPs.gen on the validated MDP.

  • mode = :branch — every call starts from the same s0. This is what a planner does when it expands one node, and it is the figure that multiplies by the per-decision node budget.
  • mode = :chain — each call continues from the previous successor, resetting to s0 on a terminal or truncated state. This is the rollout cost, and it differs from :branch because the command history and duck state evolve.
source
Duckietown.benchmark_table — Method
benchmark_table(benchmarks) -> String

Fixed-width report: per-call microseconds, allocations and kilobytes, plus the achievable call rate. Printed by the FJ8 tests and pasted into the gate document.

source
Duckietown.gen_scaling — Method
gen_scaling(mdp, s0, a; calls, mode, seed) -> Vector{GenBenchmark}

The same measurement repeated at several call counts, so a per-call figure distorted by one-off costs is visible rather than hidden.

source
Duckietown.gen_stage_profile — Method
gen_stage_profile(mdp, s0, a; calls, seed) -> Vector{GenBenchmark}

Where one gen call's time and memory actually go, measured stage by stage along the locked transition order. branch(world) is included as a reference point: it is the pure cost of deep-copying a world state, so any stage costing less than a branch is not worth optimising first.

source
Duckietown.measure — Method
measure(f, calls; label) -> GenBenchmark

Run f() calls times and record wall time, bytes allocated and allocation count. f is called once first to force compilation, and the result of every call is retained so the work cannot be optimised away. A GC.gc() runs before the measurement so the reported bytes belong to this loop.

source
Duckietown.planning_budget_estimate — Method
planning_budget_estimate(gen_bench, nodes) -> NamedTuple

What a per-decision node budget costs at the measured gen rate: seconds per decision and megabytes allocated per decision. This is the number that decides whether a planner budget is affordable, and it is derived from measurement rather than assumed.

source
Duckietown.rng_frozen — Method
rng_frozen(states, reference) -> Bool

true when every state's controller_rng still sits at exactly the stream position of reference — i.e. the shared legacy stream was never drawn from while the branches were built. This is what makes sharing it equivalent to sharing the map.

source
Duckietown.shared_by_design — Method
shared_by_design(a, b) -> Vector{String}

The objects two states are expected to share, listed explicitly so the sharing stays visible rather than accidental:

  • map and stop_signs — the read-only track description. Copying these per node would dominate a planner's memory and nothing ever writes to them.
  • controller_rng — a legacy field. Since FJ3.8 the transition's stochasticity comes from the caller's rng (x' ~ T(·|x,a) with externally supplied noise); the decision chain never draws from the state's stream. It is carried by reference because copying a MersenneTwister costs ~9 % of one gen call's total allocation for a field with no semantic role.

The sharing is only safe while the object is never written to, which is a measurable property, not an assumption — rng_frozen checks it, and world_differences compares the stream position so any advance shows up as a controller_rng mismatch.

source
Duckietown.shared_mutable_arrays — Method
shared_mutable_arrays(a, b) -> Vector{String}

Names of the per-branch mutable containers that the two states hold as the same object (===). Two branches sharing one of these would alias: writing through one would silently change the other. An empty result is the branch-purity property a planner needs.

Fields shared by design are excluded and reported separately by shared_by_design; see that docstring for why each one is safe.

source
Duckietown.world_differences — Method
world_differences(a, b) -> Vector{WorldMismatch}

Every field in which two world states differ, compared bitwise for floats and exactly for everything else, including container lengths. map is compared by identity: it is a shared read-only description of the track, not per-branch state.

source
Duckietown.BudgetPoint — Type
BudgetPoint

One (solver, budget) cell of the cost–search curve, aggregated over repeated measurements at a fixed set of states.

Latency is summarised by mean/p50/p95/max because FJ8.0 already showed wall time varies ~±20 % run to run; model_calls and the node counts are exactly reproducible, since the search is deterministic per planner seed and the repeats therefore measure timing noise only.

source
Duckietown.budget_study — Method
budget_study(mdp, make_planner, budgets; states, repeats, warmup) -> Vector{BudgetPoint}

Measure one planner family across budgets.

  • make_planner(budget, seed) returns a planner. Passing a closure is what keeps this solver-agnostic: the study never names a solver type.
  • states is the SAME state set at every budget, so the curve isolates the budget rather than mixing in state difficulty.
  • repeats re-runs each (state, budget) with the identical planner seed. The search is therefore identical and only wall time varies, which is exactly the quantity that needs repeating.
  • warmup runs one untimed decision per budget first.

mdp should be an InstrumentedMDP; otherwise model_calls is reported as -1 (unmeasured) rather than a misleading zero.

source
Duckietown.budget_table — Method
budget_table(points; extra_fields) -> String

The cost–search curve as a fixed-width table, with both budget axes side by side so an iteration count is never mistaken for a computational budget.

source
Duckietown.compute_matched_budget — Method
compute_matched_budget(points, target_calls) -> NamedTuple

The measured point whose generative-model consumption is closest to target_calls per decision, with the actual value reported rather than the target. Comparing two planners at equal iteration counts compares nothing; this is what makes an online-planner comparison fair.

Returns (; budget, actual_calls, error, latency_mean, point).

source
Duckietown.estimate_budget_for_calls — Method
estimate_budget_for_calls(points, target_calls) -> Int

The solver-native budget predicted to consume target_calls generative calls per decision, from the measured calls-per-iteration.

Selecting the nearest point of a coarse grid can be off by tens of percent; this instead inverts the measured rate, so the caller can run the estimate and report what it realised. It is an estimate from measurement, never a substitute for measuring the result.

source
Duckietown.matched_operating_points — Method
matched_operating_points(curves, targets) -> Vector{NamedTuple}

Apply compute_matched_budget to several planner curves at several generative-call targets. curves maps a label to its BudgetPoint vector. The result is the operating-point table an equal-compute comparison is run at, with each solver's realised call count recorded next to the target it was matched to.

source
Duckietown.planning_seed_config — Function
planning_seed_config(path=default) -> NamedTuple

Load the frozen planning seed split (configs/planning/seeds.yaml). The development and evaluation sets are disjoint on purpose: configuration choices are made on the former, and the six-solver comparison runs on the latter, so a planner can never be tuned against the seeds it will be reported on.

source
Duckietown.SolverRun — Type
SolverRun

Everything one solver produced under the frozen protocol: per-episode task metrics, aggregated planning cost, and the per-decision records that explain where the cost went.

source
Duckietown.SummaryStats — Type
SummaryStats

Distributional summary of one metric across episodes. n is stated because a mean over twenty episodes is not the same evidence as a mean over a thousand, and the bootstrap interval is percentile-based on a fixed RNG so the number is reproducible.

source
Duckietown.check_paired_protocol — Method
check_paired_protocol(runs) -> NamedTuple

Verify the protocol before any result is read: every solver ran the same seed set, the same number of episodes, and under the same horizon. Returns the shared seeds; throws if they differ.

source
Duckietown.cost_by_episode_position — Method
cost_by_episode_position(records; bins) -> Vector{NamedTuple}

Planning cost as a function of how far into the episode the decision was made. FJ8.4a measured a 3.8x spread in generative cost across states; a mean alone hides that, and this is what explains a high p95.

source
Duckietown.cost_table — Method
cost_table(runs) -> String

The computational-cost block. model calls = 0 means measured-and-zero: the policy performs no generative planning. n/a means unmeasured.

source
Duckietown.episode_csv — Method
episode_csv(runs) -> String

Episode-level records for every solver, one row per (solver, seed), so the paired structure survives outside this session and the comparison can be reanalysed without re-running it.

source
Duckietown.paired_difference — Method
paired_difference(name_a, eps_a, name_b, eps_b, metric) -> PairedDifference

Per-seed differences on a shared seed set. Throws if the two solvers did not run the same seeds — silently intersecting them would turn a protocol error into a quieter, wrong answer.

source
Duckietown.safety_table — Method
safety_table(runs) -> String

Counts, with the episode total as the explicit denominator, plus stop compliance over its own denominator.

source
Duckietown.stop_compliance — Method
stop_compliance(episodes) -> Union{Nothing,Float64}

Compliant stop encounters over stop encounters, where an encounter is a passed_stop event — a stop sign the agent actually traversed. nothing when no stop sign was ever reached, so an untested solver is never credited with perfect compliance.

source
Duckietown.summary_stats — Method
summary_stats(values; bootstrap, seed) -> SummaryStats

Mean, median, SD, IQR, range and a percentile bootstrap 95 % CI of the mean.

source
Duckietown.task_table — Method
task_table(runs) -> String

The task-performance block. Contains no timing and no model-call column by construction: computational cost lives in cost_table.

source
Duckietown.DecisionRecord — Type
DecisionRecord

One decision's cost, with the episode and step it belongs to. FJ8.4a measured a 3.8x spread in generative cost across states, so a flat list of diagnostics throws away exactly the structure needed to explain a high p95 latency.

source
Duckietown.DecisionTrace — Type
DecisionTrace

One decision, recorded observationally. Every field is a value the evaluator had in hand at the moment the decision completed.

source
Duckietown.EpisodeMetrics — Type
EpisodeMetrics

Outcome of one evaluated episode. Every field is derived from the rollout log, never from a separate bookkeeping path.

source
Duckietown.PlannerCost — Type
PlannerCost

What producing the decisions cost, aggregated over an evaluation. This is never combined with return, progress or any other task metric: a slow planner that drives well and a fast one that drives badly must stay distinguishable.

extra holds the mean of every numeric field the planner reported in its PlanningDiagnostics.extra, so a tree search's node counts and a particle method's belief counts both aggregate without the evaluator knowing either.

source
Duckietown._trace — Method
_trace(solver, seed, k, s, a, r, diag, mdp) -> DecisionTrace

Assemble one row from values already computed. No call here recomputes anything the decision did not already produce.

source
Duckietown.compare_policies — Method
compare_policies(mdp_by_name, policies; seeds, max_steps) -> Dict

Evaluate several policies and return one summary per name. Each policy is run on the MDP appropriate to its action space (tabular policies need the discrete variant, actor policies the continuous one), which is why the MDPs are passed per name — the problem definition is the same either way, only the action representation differs.

source
Duckietown.evaluate_planner — Method
evaluate_planner(mdp, planner; seeds, max_steps) -> (episodes, cost)

Evaluate any planner through the same evaluate_policy used for the learned and tabular baselines, and return its planning cost alongside — two separate values, deliberately not one score.

Wrap mdp in an InstrumentedMDP before solve to have model_calls measured; without it the cost report says so rather than claiming zero.

source
Duckietown.evaluate_policy — Method
evaluate_policy(mdp, policy; seeds, max_steps, stop_zone, yield_speed)
    -> Vector{EpisodeMetrics}

Run one episode per seed: sample x0 from initialstate(mdp), then act greedily with policy until the episode ends or max_steps decisions elapse. Returns the per-episode metrics; use summarize_evaluation to aggregate.

Pass record a Vector{PlanningDiagnostics} — or a Vector{DecisionRecord} to keep the episode and step each measurement belongs to — to also collect what each decision cost. This changes nothing about the episode — the same action is taken either way — it only routes the call through plan_action so the cost can be observed. Task performance and planning cost are reported separately and never mixed into one score.

source
Duckietown.reaggregate_episodes — Method
reaggregate_episodes(traces) -> Vector{EpisodeMetrics}

Rebuild episode metrics from the per-decision trace, using the same definitions evaluate_policy uses.

If the two ever disagree, that is a genuine finding about the enrichment run and not something to reconcile with a tolerance — the protocol is frozen and deterministic, so exact reproduction is the expectation.

source
Duckietown.summarize_planning — Method
summarize_planning(diagnostics) -> PlannerCost

Aggregate per-decision diagnostics. model_calls_total is -1 when the model was not instrumented, rather than 0, so "not measured" never reads as "free".

source
Duckietown.LIBM_DERIVED_FIELDS — Constant
LIBM_DERIVED_FIELDS

The only quantities allowed to differ at all, and the reason each may.

Root cause (measured, not assumed). The dynamical state itself is bit-identical: after a matched-state step the DB18 pose/velocity matrices q0/v0 agree bit-for-bit. The two runtimes diverge only where they read that state back through their own libm. Recomputing atan2(q0[2,1], q0[1,1]) in Julia on the reference's own q0 reproduces the Julia angle exactly while the reference's stored angle differs by 1 ULP, so the source is atan2 (OpenLibm vs glibc) — the deviation class already recorded in FJ2/FJ3, now confirmed live.

Propagation chain (measured worst case, Julia 1.11.3 vs ddm-ref):

ego.angle          1 ULP   atan2 pose readback              <- root
  lane_fallback[2] 2 ULP   acos(dot(dir(angle), tangent)), ill-conditioned
    raw.phi        2 ULP   the clamped lane angle
      reward.heading 4 ULP  -alpha*phi^2 doubles the relative error
      reward.total   4 ULP  inherits the heading term
  reward.progress  1 ULP   alpha*v*cos(phi)

Worst absolute difference over a 40-decision matched-state sweep: 2.22e-16 — twelve orders of magnitude below the quantities involved (rewards ~1e-1 rad/units). raw.d, the ego pose, speed, the duckie state, the delay window, the events and the termination classification are all bit-identical.

Second, independent libm source (found by FJ6). The quadratic reward terms are written state.phi ** 2 / state.d ** 2 in the reference, and CPython's float.__pow__ is a libm pow() call — not a multiplication. Measured at phi = 0.16103364894924665:

Python  phi ** 2   = 0.02593183609390921    (glibc pow)
Python  phi * phi  = 0.025931836093909207
Julia   phi^2      = 0.025931836093909207   (x*x)
Julia   phi^2.0    = 0.025931836093909207   (OpenLibm pow agrees with x*x)

so glibc's pow is 1 ULP off the correctly-rounded square here. Julia cannot reproduce that without emulating glibc's pow, which would make the model platform-dependent — a worse outcome than a 1-ULP reward difference that never touches the state. reward.lateral is included for the same reason (d ** 2). Unlike the atan2 source, this one is NOT affected by in-process interposition: embedded Python still returned glibc's value.

The exact set is toolchain-dependent: on Julia 1.10.11 only ego.angle deviated (1 ULP); 1.11.3 propagates it a little further. The invariant that matters — and that the tests assert — is STRUCTURAL: nothing outside this libm-derived chain may deviate at all, and no discrete field may disagree.

source
Duckietown.LIBM_MAX_ULPS — Constant
LIBM_MAX_ULPS / LIBM_MAX_ABSDIFF

Numeric bounds for [LIBM_DERIVED_FIELDS], derived from measurement rather than chosen (master prompt §33): the observed worst case is 4 ULP and 2.22e-16 absolute; the bounds keep 2x ULP headroom for other toolchains while staying far below any physically meaningful scale.

source
Duckietown.SIGNED_ZERO_FIELDS — Constant
SIGNED_ZERO_FIELDS

Entries that can carry a different ZERO SIGN between the runtimes with no numerical difference whatsoever (0.0 == -0.0).

Measured: the se(2) body-velocity diagonal ego.v0[1,1] / ego.v0[2,2] on turning actions (the reference builds the matrix through NumPy products that can yield -0.0). These entries are never read: the only consumers of v0 are linear_angular_from_se2 (which reads [1,3], [2,3], [2,1]) and SE2_from_se2 (which reads [2,1] and [1:2,3]). The diagonal is inert, so this cannot affect any transition.

source
Duckietown.FieldDiff — Type
FieldDiff

One compared quantity: its name, the two values, the ULP distance (for finite Float64 pairs) and the absolute difference. ulps == 0 means bit-identical.

source
Duckietown.StepParityReport — Type
StepParityReport

Result of one matched-state comparison: the field-level diffs, the worst ULP/absolute distances, and the discrete agreements (reason, terminated, truncated, events, tile, duck class, counters). exact means every compared quantity was bit-identical and every discrete field agreed.

source
Duckietown._ulps — Method
_ulps(a, b) -> Int

ULP distance under the IEEE-754 total order, saturating instead of overflowing. Numerically equal values are 0 ULP apart — including 0.0 vs -0.0, whose raw reinterpret difference is typemin(Int64) and overflowed an earlier abs(...)-based implementation into a negative "distance" (a measurement-tool bug, caught by FJ5.3 refusing to accept otherwise clean steps).

source
Duckietown.bitwise_only_fields — Method
bitwise_only_fields(reports) -> Vector{String}

Fields that are numerically equal but differ in bit pattern — in practice the signed-zero entries described in [SIGNED_ZERO_FIELDS].

source
Duckietown.compare_worlds — Method
compare_worlds(julia_world, reference_world) -> Vector{FieldDiff}

Field-level comparison of the full latent state: ego pose/speed/DB18 matrices/wheel axes/delay window, every duckie's object state, and the memories. Discrete fields are checked by compare_step.

source
Duckietown.nonzero_fields — Method
nonzero_fields(reports) -> Vector{String}

Every field name that was ever non-bit-identical across a sweep. The FJ5 evidence is that this set equals [LIBM_1ULP_FIELDS] or is empty.

source
Duckietown.parity_accepted — Method
parity_accepted(report; max_libm_ulps=LIBM_MAX_ULPS,
                max_libm_absdiff=LIBM_MAX_ABSDIFF) -> Bool

The FJ5 acceptance criterion: every discrete field agrees and every compared quantity is bit-identical, EXCEPT the attributed [LIBM_DERIVED_FIELDS], which may differ within the measured bounds.

source
Duckietown.parity_summary — Method
parity_summary(reports) -> NamedTuple

Aggregate a sweep: number of steps, how many were bit-exact, the worst ULP and absolute distances with the field that produced them, and every discrete mismatch seen.

source
Duckietown.DriftReport — Type
DriftReport

Per-decision drift between two free-running rollouts, plus the milestone decisions that matter more than any aggregate error.

source
Duckietown.RolloutRecord — Type
RolloutRecord

One decision of a free-running rollout. Everything needed to compare two runs without re-deriving anything: the dynamical state (q0, v0, delay window length), the pose readback, the tabular projection, the stop and duck state, the action actually applied, the full reward breakdown and the episode flags.

source
Duckietown.compare_rollouts — Method
compare_rollouts(a, b) -> DriftReport

Compare two free-running rollouts decision by decision. a is the reference run, b the run under test (native Julia, or the other transport).

source
Duckietown.drift_summary — Method
drift_summary(report) -> NamedTuple

Compact, JSON-ready summary: milestone decisions, final and maximum drifts, and the Type-1/Type-2 classification.

source
Duckietown.event_timing — Method
event_timing(records) -> Dict{String,Union{Nothing,Int}}

Decision index at which each tracked event/condition first occurs in a rollout (nothing if it never does).

source
Duckietown.libm_hypothesis_check — Method
libm_hypothesis_check(rows) -> NamedTuple

Aggregate the three-lane table into the FJ5-R prediction test: Δ_PJ ≈ Δ_PP and Δ_CJ ≈ 0 for every libm-sensitive field.

d_CJ_exactly_zero is the strict form. It does NOT always hold, and that is itself informative: the in-process interposition makes CPython's math.atan2 resolve to Julia's libm, but NumPy's ufunc path has its own inner loop, so for occasional inputs the embedded reference still lands 1 ULP away from Julia. max_abs_d_CJ therefore reports the measured magnitude instead of hiding it behind a boolean.

source
Duckietown.rollout_native — Method
rollout_native(model, x0, actions; rng, discount=1.0) -> Vector{RolloutRecord}

Free-running native rollout: simulate_decision applied repeatedly, feeding its own successor state back in. Stops early on a genuine terminal.

source
Duckietown.three_lane_table — Method
three_lane_table(process, pycall, julia, fields) -> Vector{Dict}

The FJ5-R-motivated three-value log: for each decision and each libm-sensitive field, the value from the isolated Python reference, the in-process Python oracle and native Julia, plus

Δ_PJ = process - julia      (true cross-runtime difference)
Δ_PP = process - pycall     (libm isolation effect)
Δ_CJ = pycall  - julia      (expected ~0 if the libm hypothesis holds)
source