Evaluation
Parity, rollouts, metrics, budget and comparison harnesses.
Duckietown.GenBenchmark — Type
GenBenchmarkMeasured cost of calls repetitions of one operation. Per-call figures come from per_call_us, bytes_per_call and allocs_per_call.
Duckietown.WorldMismatch — Type
WorldMismatchOne field where two world states differ, with both values rendered for the report. Floats are compared bitwise (===), so 0.0 and -0.0 count as different and NaN equals NaN — a planner-purity check must not be fooled by numeric coincidence.
Duckietown.benchmark_gen — Method
benchmark_gen(mdp, s0, a; calls, mode, seed) -> GenBenchmarkCost of POMDPs.gen on the validated MDP.
mode = :branch— every call starts from the sames0. This is what a planner does when it expands one node, and it is the figure that multiplies by the per-decision node budget.mode = :chain— each call continues from the previous successor, resetting tos0on a terminal or truncated state. This is the rollout cost, and it differs from:branchbecause the command history and duck state evolve.
Duckietown.benchmark_table — Method
benchmark_table(benchmarks) -> StringFixed-width report: per-call microseconds, allocations and kilobytes, plus the achievable call rate. Printed by the FJ8 tests and pasted into the gate document.
Duckietown.gen_scaling — Method
gen_scaling(mdp, s0, a; calls, mode, seed) -> Vector{GenBenchmark}The same measurement repeated at several call counts, so a per-call figure distorted by one-off costs is visible rather than hidden.
Duckietown.gen_stage_profile — Method
gen_stage_profile(mdp, s0, a; calls, seed) -> Vector{GenBenchmark}Where one gen call's time and memory actually go, measured stage by stage along the locked transition order. branch(world) is included as a reference point: it is the pure cost of deep-copying a world state, so any stage costing less than a branch is not worth optimising first.
Duckietown.measure — Method
measure(f, calls; label) -> GenBenchmarkRun f() calls times and record wall time, bytes allocated and allocation count. f is called once first to force compilation, and the result of every call is retained so the work cannot be optimised away. A GC.gc() runs before the measurement so the reported bytes belong to this loop.
Duckietown.planning_budget_estimate — Method
planning_budget_estimate(gen_bench, nodes) -> NamedTupleWhat a per-decision node budget costs at the measured gen rate: seconds per decision and megabytes allocated per decision. This is the number that decides whether a planner budget is affordable, and it is derived from measurement rather than assumed.
Duckietown.rng_frozen — Method
rng_frozen(states, reference) -> Booltrue when every state's controller_rng still sits at exactly the stream position of reference — i.e. the shared legacy stream was never drawn from while the branches were built. This is what makes sharing it equivalent to sharing the map.
Duckietown.shared_by_design — Method
shared_by_design(a, b) -> Vector{String}The objects two states are expected to share, listed explicitly so the sharing stays visible rather than accidental:
mapandstop_signs— the read-only track description. Copying these per node would dominate a planner's memory and nothing ever writes to them.controller_rng— a legacy field. Since FJ3.8 the transition's stochasticity comes from the caller'srng(x' ~ T(·|x,a)with externally supplied noise); the decision chain never draws from the state's stream. It is carried by reference because copying aMersenneTwistercosts ~9 % of onegencall's total allocation for a field with no semantic role.
The sharing is only safe while the object is never written to, which is a measurable property, not an assumption — rng_frozen checks it, and world_differences compares the stream position so any advance shows up as a controller_rng mismatch.
Duckietown.shared_mutable_arrays — Method
shared_mutable_arrays(a, b) -> Vector{String}Names of the per-branch mutable containers that the two states hold as the same object (===). Two branches sharing one of these would alias: writing through one would silently change the other. An empty result is the branch-purity property a planner needs.
Fields shared by design are excluded and reported separately by shared_by_design; see that docstring for why each one is safe.
Duckietown.world_differences — Method
world_differences(a, b) -> Vector{WorldMismatch}Every field in which two world states differ, compared bitwise for floats and exactly for everything else, including container lengths. map is compared by identity: it is a shared read-only description of the track, not per-branch state.
Duckietown.worlds_identical — Method
worlds_identical(a, b) -> Booltrue when world_differences is empty.
Duckietown.BudgetPoint — Type
BudgetPointOne (solver, budget) cell of the cost–search curve, aggregated over repeated measurements at a fixed set of states.
Latency is summarised by mean/p50/p95/max because FJ8.0 already showed wall time varies ~±20 % run to run; model_calls and the node counts are exactly reproducible, since the search is deterministic per planner seed and the repeats therefore measure timing noise only.
Duckietown.budget_study — Method
budget_study(mdp, make_planner, budgets; states, repeats, warmup) -> Vector{BudgetPoint}Measure one planner family across budgets.
make_planner(budget, seed)returns a planner. Passing a closure is what keeps this solver-agnostic: the study never names a solver type.statesis the SAME state set at every budget, so the curve isolates the budget rather than mixing in state difficulty.repeatsre-runs each (state, budget) with the identical planner seed. The search is therefore identical and only wall time varies, which is exactly the quantity that needs repeating.warmupruns one untimed decision per budget first.
mdp should be an InstrumentedMDP; otherwise model_calls is reported as -1 (unmeasured) rather than a misleading zero.
Duckietown.budget_table — Method
budget_table(points; extra_fields) -> StringThe cost–search curve as a fixed-width table, with both budget axes side by side so an iteration count is never mistaken for a computational budget.
Duckietown.compute_matched_budget — Method
compute_matched_budget(points, target_calls) -> NamedTupleThe measured point whose generative-model consumption is closest to target_calls per decision, with the actual value reported rather than the target. Comparing two planners at equal iteration counts compares nothing; this is what makes an online-planner comparison fair.
Returns (; budget, actual_calls, error, latency_mean, point).
Duckietown.estimate_budget_for_calls — Method
estimate_budget_for_calls(points, target_calls) -> IntThe solver-native budget predicted to consume target_calls generative calls per decision, from the measured calls-per-iteration.
Selecting the nearest point of a coarse grid can be off by tens of percent; this instead inverts the measured rate, so the caller can run the estimate and report what it realised. It is an estimate from measurement, never a substitute for measuring the result.
Duckietown.matched_operating_points — Method
matched_operating_points(curves, targets) -> Vector{NamedTuple}Apply compute_matched_budget to several planner curves at several generative-call targets. curves maps a label to its BudgetPoint vector. The result is the operating-point table an equal-compute comparison is run at, with each solver's realised call count recorded next to the target it was matched to.
Duckietown.operating_point_table — Method
operating_point_table(rows) -> StringDuckietown.planning_seed_config — Function
planning_seed_config(path=default) -> NamedTupleLoad the frozen planning seed split (configs/planning/seeds.yaml). The development and evaluation sets are disjoint on purpose: configuration choices are made on the former, and the six-solver comparison runs on the latter, so a planner can never be tuned against the seeds it will be reported on.
Duckietown.PairedDifference — Type
PairedDifferencea − b on the episodes both solvers ran, matched by seed.
Duckietown.SolverRun — Type
SolverRunEverything one solver produced under the frozen protocol: per-episode task metrics, aggregated planning cost, and the per-decision records that explain where the cost went.
Duckietown.SummaryStats — Type
SummaryStatsDistributional summary of one metric across episodes. n is stated because a mean over twenty episodes is not the same evidence as a mean over a thousand, and the bootstrap interval is percentile-based on a fixed RNG so the number is reproducible.
Duckietown.check_paired_protocol — Method
check_paired_protocol(runs) -> NamedTupleVerify the protocol before any result is read: every solver ran the same seed set, the same number of episodes, and under the same horizon. Returns the shared seeds; throws if they differ.
Duckietown.cost_by_episode_position — Method
cost_by_episode_position(records; bins) -> Vector{NamedTuple}Planning cost as a function of how far into the episode the decision was made. FJ8.4a measured a 3.8x spread in generative cost across states; a mean alone hides that, and this is what explains a high p95.
Duckietown.cost_table — Method
cost_table(runs) -> StringThe computational-cost block. model calls = 0 means measured-and-zero: the policy performs no generative planning. n/a means unmeasured.
Duckietown.episode_csv — Method
episode_csv(runs) -> StringEpisode-level records for every solver, one row per (solver, seed), so the paired structure survives outside this session and the comparison can be reanalysed without re-running it.
Duckietown.paired_difference — Method
paired_difference(name_a, eps_a, name_b, eps_b, metric) -> PairedDifferencePer-seed differences on a shared seed set. Throws if the two solvers did not run the same seeds — silently intersecting them would turn a protocol error into a quieter, wrong answer.
Duckietown.paired_table — Method
paired_table(diffs) -> StringDuckietown.position_table — Method
position_table(rows) -> StringDuckietown.safety_table — Method
safety_table(runs) -> StringCounts, with the episode total as the explicit denominator, plus stop compliance over its own denominator.
Duckietown.stop_compliance — Method
stop_compliance(episodes) -> Union{Nothing,Float64}Compliant stop encounters over stop encounters, where an encounter is a passed_stop event — a stop sign the agent actually traversed. nothing when no stop sign was ever reached, so an untested solver is never credited with perfect compliance.
Duckietown.summary_stats — Method
summary_stats(values; bootstrap, seed) -> SummaryStatsMean, median, SD, IQR, range and a percentile bootstrap 95 % CI of the mean.
Duckietown.task_table — Method
task_table(runs) -> StringThe task-performance block. Contains no timing and no model-call column by construction: computational cost lives in cost_table.
Duckietown.DecisionRecord — Type
DecisionRecordOne decision's cost, with the episode and step it belongs to. FJ8.4a measured a 3.8x spread in generative cost across states, so a flat list of diagnostics throws away exactly the structure needed to explain a high p95 latency.
Duckietown.DecisionTrace — Type
DecisionTraceOne decision, recorded observationally. Every field is a value the evaluator had in hand at the moment the decision completed.
Duckietown.EpisodeMetrics — Type
EpisodeMetricsOutcome of one evaluated episode. Every field is derived from the rollout log, never from a separate bookkeeping path.
Duckietown.PlannerCost — Type
PlannerCostWhat producing the decisions cost, aggregated over an evaluation. This is never combined with return, progress or any other task metric: a slow planner that drives well and a fast one that drives badly must stay distinguishable.
extra holds the mean of every numeric field the planner reported in its PlanningDiagnostics.extra, so a tree search's node counts and a particle method's belief counts both aggregate without the evaluator knowing either.
Duckietown._trace — Method
_trace(solver, seed, k, s, a, r, diag, mdp) -> DecisionTraceAssemble one row from values already computed. No call here recomputes anything the decision did not already produce.
Duckietown.compare_policies — Method
compare_policies(mdp_by_name, policies; seeds, max_steps) -> DictEvaluate several policies and return one summary per name. Each policy is run on the MDP appropriate to its action space (tabular policies need the discrete variant, actor policies the continuous one), which is why the MDPs are passed per name — the problem definition is the same either way, only the action representation differs.
Duckietown.decision_csv — Method
decision_csv(traces) -> StringDuckietown.evaluate_planner — Method
evaluate_planner(mdp, planner; seeds, max_steps) -> (episodes, cost)Evaluate any planner through the same evaluate_policy used for the learned and tabular baselines, and return its planning cost alongside — two separate values, deliberately not one score.
Wrap mdp in an InstrumentedMDP before solve to have model_calls measured; without it the cost report says so rather than claiming zero.
Duckietown.evaluate_policy — Method
evaluate_policy(mdp, policy; seeds, max_steps, stop_zone, yield_speed)
-> Vector{EpisodeMetrics}Run one episode per seed: sample x0 from initialstate(mdp), then act greedily with policy until the episode ends or max_steps decisions elapse. Returns the per-episode metrics; use summarize_evaluation to aggregate.
Pass record a Vector{PlanningDiagnostics} — or a Vector{DecisionRecord} to keep the episode and step each measurement belongs to — to also collect what each decision cost. This changes nothing about the episode — the same action is taken either way — it only routes the call through plan_action so the cost can be observed. Task performance and planning cost are reported separately and never mixed into one score.
Duckietown.reaggregate_episodes — Method
reaggregate_episodes(traces) -> Vector{EpisodeMetrics}Rebuild episode metrics from the per-decision trace, using the same definitions evaluate_policy uses.
If the two ever disagree, that is a genuine finding about the enrichment run and not something to reconcile with a tolerance — the protocol is frozen and deterministic, so exact reproduction is the expectation.
Duckietown.summarize_evaluation — Method
summarize_evaluation(metrics) -> NamedTupleAggregate episodes into the comparison row used across solvers.
Duckietown.summarize_planning — Method
summarize_planning(diagnostics) -> PlannerCostAggregate per-decision diagnostics. model_calls_total is -1 when the model was not instrumented, rather than 0, so "not measured" never reads as "free".
Duckietown.LIBM_DERIVED_FIELDS — Constant
LIBM_DERIVED_FIELDSThe only quantities allowed to differ at all, and the reason each may.
Root cause (measured, not assumed). The dynamical state itself is bit-identical: after a matched-state step the DB18 pose/velocity matrices q0/v0 agree bit-for-bit. The two runtimes diverge only where they read that state back through their own libm. Recomputing atan2(q0[2,1], q0[1,1]) in Julia on the reference's own q0 reproduces the Julia angle exactly while the reference's stored angle differs by 1 ULP, so the source is atan2 (OpenLibm vs glibc) — the deviation class already recorded in FJ2/FJ3, now confirmed live.
Propagation chain (measured worst case, Julia 1.11.3 vs ddm-ref):
ego.angle 1 ULP atan2 pose readback <- root
lane_fallback[2] 2 ULP acos(dot(dir(angle), tangent)), ill-conditioned
raw.phi 2 ULP the clamped lane angle
reward.heading 4 ULP -alpha*phi^2 doubles the relative error
reward.total 4 ULP inherits the heading term
reward.progress 1 ULP alpha*v*cos(phi)Worst absolute difference over a 40-decision matched-state sweep: 2.22e-16 — twelve orders of magnitude below the quantities involved (rewards ~1e-1 rad/units). raw.d, the ego pose, speed, the duckie state, the delay window, the events and the termination classification are all bit-identical.
Second, independent libm source (found by FJ6). The quadratic reward terms are written state.phi ** 2 / state.d ** 2 in the reference, and CPython's float.__pow__ is a libm pow() call — not a multiplication. Measured at phi = 0.16103364894924665:
Python phi ** 2 = 0.02593183609390921 (glibc pow)
Python phi * phi = 0.025931836093909207
Julia phi^2 = 0.025931836093909207 (x*x)
Julia phi^2.0 = 0.025931836093909207 (OpenLibm pow agrees with x*x)so glibc's pow is 1 ULP off the correctly-rounded square here. Julia cannot reproduce that without emulating glibc's pow, which would make the model platform-dependent — a worse outcome than a 1-ULP reward difference that never touches the state. reward.lateral is included for the same reason (d ** 2). Unlike the atan2 source, this one is NOT affected by in-process interposition: embedded Python still returned glibc's value.
The exact set is toolchain-dependent: on Julia 1.10.11 only ego.angle deviated (1 ULP); 1.11.3 propagates it a little further. The invariant that matters — and that the tests assert — is STRUCTURAL: nothing outside this libm-derived chain may deviate at all, and no discrete field may disagree.
Duckietown.LIBM_MAX_ULPS — Constant
LIBM_MAX_ULPS / LIBM_MAX_ABSDIFFNumeric bounds for [LIBM_DERIVED_FIELDS], derived from measurement rather than chosen (master prompt §33): the observed worst case is 4 ULP and 2.22e-16 absolute; the bounds keep 2x ULP headroom for other toolchains while staying far below any physically meaningful scale.
Duckietown.SIGNED_ZERO_FIELDS — Constant
SIGNED_ZERO_FIELDSEntries that can carry a different ZERO SIGN between the runtimes with no numerical difference whatsoever (0.0 == -0.0).
Measured: the se(2) body-velocity diagonal ego.v0[1,1] / ego.v0[2,2] on turning actions (the reference builds the matrix through NumPy products that can yield -0.0). These entries are never read: the only consumers of v0 are linear_angular_from_se2 (which reads [1,3], [2,3], [2,1]) and SE2_from_se2 (which reads [2,1] and [1:2,3]). The diagonal is inert, so this cannot affect any transition.
Duckietown.FieldDiff — Type
FieldDiffOne compared quantity: its name, the two values, the ULP distance (for finite Float64 pairs) and the absolute difference. ulps == 0 means bit-identical.
Duckietown.StepParityReport — Type
StepParityReportResult of one matched-state comparison: the field-level diffs, the worst ULP/absolute distances, and the discrete agreements (reason, terminated, truncated, events, tile, duck class, counters). exact means every compared quantity was bit-identical and every discrete field agreed.
Duckietown._ulps — Method
_ulps(a, b) -> IntULP distance under the IEEE-754 total order, saturating instead of overflowing. Numerically equal values are 0 ULP apart — including 0.0 vs -0.0, whose raw reinterpret difference is typemin(Int64) and overflowed an earlier abs(...)-based implementation into a negative "distance" (a measurement-tool bug, caught by FJ5.3 refusing to accept otherwise clean steps).
Duckietown.bitwise_only_fields — Method
bitwise_only_fields(reports) -> Vector{String}Fields that are numerically equal but differ in bit pattern — in practice the signed-zero entries described in [SIGNED_ZERO_FIELDS].
Duckietown.compare_worlds — Method
compare_worlds(julia_world, reference_world) -> Vector{FieldDiff}Field-level comparison of the full latent state: ego pose/speed/DB18 matrices/wheel axes/delay window, every duckie's object state, and the memories. Discrete fields are checked by compare_step.
Duckietown.nonzero_fields — Method
nonzero_fields(reports) -> Vector{String}Every field name that was ever non-bit-identical across a sweep. The FJ5 evidence is that this set equals [LIBM_1ULP_FIELDS] or is empty.
Duckietown.parity_accepted — Method
parity_accepted(report; max_libm_ulps=LIBM_MAX_ULPS,
max_libm_absdiff=LIBM_MAX_ABSDIFF) -> BoolThe FJ5 acceptance criterion: every discrete field agrees and every compared quantity is bit-identical, EXCEPT the attributed [LIBM_DERIVED_FIELDS], which may differ within the measured bounds.
Duckietown.parity_summary — Method
parity_summary(reports) -> NamedTupleAggregate a sweep: number of steps, how many were bit-exact, the worst ULP and absolute distances with the field that produced them, and every discrete mismatch seen.
Duckietown.DriftReport — Type
DriftReportPer-decision drift between two free-running rollouts, plus the milestone decisions that matter more than any aggregate error.
Duckietown.RolloutRecord — Type
RolloutRecordOne decision of a free-running rollout. Everything needed to compare two runs without re-deriving anything: the dynamical state (q0, v0, delay window length), the pose readback, the tabular projection, the stop and duck state, the action actually applied, the full reward breakdown and the episode flags.
Duckietown.compare_rollouts — Method
compare_rollouts(a, b) -> DriftReportCompare two free-running rollouts decision by decision. a is the reference run, b the run under test (native Julia, or the other transport).
Duckietown.drift_summary — Method
drift_summary(report) -> NamedTupleCompact, JSON-ready summary: milestone decisions, final and maximum drifts, and the Type-1/Type-2 classification.
Duckietown.event_timing — Method
event_timing(records) -> Dict{String,Union{Nothing,Int}}Decision index at which each tracked event/condition first occurs in a rollout (nothing if it never does).
Duckietown.event_timing_diff — Method
event_timing_diff(a, b) -> Dict{String,Any}Per-event first-occurrence indices in both runs and their difference.
Duckietown.libm_hypothesis_check — Method
libm_hypothesis_check(rows) -> NamedTupleAggregate the three-lane table into the FJ5-R prediction test: Δ_PJ ≈ Δ_PP and Δ_CJ ≈ 0 for every libm-sensitive field.
d_CJ_exactly_zero is the strict form. It does NOT always hold, and that is itself informative: the in-process interposition makes CPython's math.atan2 resolve to Julia's libm, but NumPy's ufunc path has its own inner loop, so for occasional inputs the embedded reference still lands 1 ULP away from Julia. max_abs_d_CJ therefore reports the measured magnitude instead of hiding it behind a boolean.
Duckietown.rollout_native — Method
rollout_native(model, x0, actions; rng, discount=1.0) -> Vector{RolloutRecord}Free-running native rollout: simulate_decision applied repeatedly, feeding its own successor state back in. Stops early on a genuine terminal.
Duckietown.rollout_table — Method
rollout_table(records) -> StringCSV text of a rollout log (one row per decision), for the FJ6 artifacts.
Duckietown.three_lane_table — Method
three_lane_table(process, pycall, julia, fields) -> Vector{Dict}The FJ5-R-motivated three-value log: for each decision and each libm-sensitive field, the value from the isolated Python reference, the in-process Python oracle and native Julia, plus
Δ_PJ = process - julia (true cross-runtime difference)
Δ_PP = process - pycall (libm isolation effect)
Δ_CJ = pycall - julia (expected ~0 if the libm hypothesis holds)