Interfaces
POMDPs.jl, the RL-style environment, planning contract, reproducibility.
Duckietown.AnyMDPLike — Type
AnyMDPLikeMDPLike for either action type.
Duckietown.MDPLike — Type
MDPLike{A}The model or a transparent wrapper around it, for a given action type. Code that only reads the model accepts this so a diagnostic wrapper never forces a second implementation of anything.
Duckietown.InstrumentedMDP — Type
InstrumentedMDP(mdp)The same model, counting generative calls. It forwards every POMDPs.jl function and — through getproperty — every field, so anything written against DuckietownMDP works on it unchanged.
It is a measuring device, not a translation layer: it must not, and does not, alter states, rewards, transitions, action bounds or termination. That transparency is asserted bitwise in test/test_fj8_planning.jl rather than merely claimed here.
Only calls through POMDPs.gen are counted, which is exactly a planner's consumption of the model — the evaluator's own environment step goes through simulate_decision and is deliberately not charged as planning cost.
Duckietown.PlanningDiagnostics — Type
PlanningDiagnostics(planning_time, model_calls, extra)What producing one action cost. Only two fields are universal:
planning_time— wall-clock seconds for the decision;model_calls— generative-model calls consumed, or-1when not measured.
Everything solver-specific goes in extra, an open NamedTuple. A tree search may report (tree_nodes = ..., max_depth = ..., iterations = ...); a value-iteration style solver (backups = ..., convergence_iterations = ...); a particle method (particles = ..., belief_nodes = ...). The evaluator never looks inside, so it cannot assume a planner has a tree — or has any internal structure at all, which is the normal case for a learned policy.
Duckietown._probe_states — Method
States along one driven trajectory, with the action taken at each, used so capability probing is not confined to the spawn state. Returns (states, actions).
Duckietown.capability_report — Method
capability_report(mdp) -> Stringmodel_capabilities rendered as the compatibility table used in the gate documents.
Duckietown.model_calls — Method
model_calls(mdp) -> IntGenerative calls consumed so far; -1 for an uninstrumented model, so callers report "not measured" rather than a misleading zero.
Duckietown.model_capabilities — Method
model_capabilities(mdp; seeds) -> NamedTupleWhat this model actually offers, determined by exercising the interface rather than by asserting it. The question a new solver should raise is "are this solver's requirements met by the model?", not "can the model be bent to fit the solver?" — this is the answer to the first form.
Stochasticity is reported as three separate facts, because this model is conditionally stochastic and collapsing that into one flag would mislead a planner:
consumes_rng— the transition draws from the caller's stream at some reachable state;stochastic_state_fraction— at what fraction of the probed states it does so. The pedestrian trigger only fires when a duck is armed, ahead and inside the trigger distance window, so most states have a deterministic successor;stochastic_outcomes— whether two seeds were actually observed to produce different successors.
A planner that widens over states (DPW's k_state/alpha_state) is doing useful work only in the fraction of states counted here; elsewhere the successor is a function of (s, a) alone. A false is a statement about the states and seeds probed, never a proof of determinism.
Pass policy to drive the probe trajectory with a competent controller. This matters: under a constant action the vehicle leaves the road within a few decisions and never reaches the states where a pedestrian can trigger, so the default probe under-reports what the model can do. The policy is only a way of reaching representative states — no property reported here depends on which policy is used, only on which states it visits.
Duckietown.plan_action — Method
plan_action(policy, mdp, s) -> (action, PlanningDiagnostics)One decision plus what it cost. The default times policy_action and reads the model's call counter, which already works for any solver without the solver knowing this package exists.
A solver extension may add a method that fills richer extra fields. That is the only thing an extension is ever expected to add, and it is optional.
Duckietown.policy_action — Method
policy_action(policy, mdp, s) -> actionAsk any policy for its action, reconciling the two calling conventions in play — inside this package, without adding a method to POMDPs.action.
POMDPs.jl's contract is action(policy, x): solve has already bound the model into the policy, so a planner needs no model argument. This package's evaluator passes the model explicitly, because a stateless policy (a Q-table, an actor network) is reusable across models and needs it.
Extending POMDPs.action with a three-argument form would change how a generic POMDPs.jl policy behaves for everyone who loads this package — the opposite of the goal, which is that solvers plug in without the model altering the ecosystem around them. So the adaptation lives here, on a function this package owns:
- a
POMDPs.Policy(anythingsolvereturns) is asked the standard way; - an
AbstractPolicy(this package's tabular and actor adapters) is asked with the model, which is how those are defined; - anything else falls back to the three-argument form, so a hand-written policy that defines it keeps working.
Duckietown.reset_model_calls! — Method
reset_model_calls!(mdp) -> IntZero the counter and return its previous value.
Duckietown.AbstractPolicy — Type
AbstractPolicyInterface boundary for reference-policy adapters (solvers/adapters.jl) and future solver policies. An adapter wraps one shipped artefact (policies/*/policy.npy or policy.pt) and maps the appropriate projection (discrete or encoded continuous observation) to an action.
Duckietown.act — Function
act(policy, observation, rng) -> actionReturn the policy's action for the given observation (post-encoding continuous vector for SAC/TD3 adapters, raw state for tabular adapters). rng is required for stochastic exploration behaviour; deterministic evaluation must not consume it.
Duckietown.VISUALIZATION_EXTENSION_POINTS — Constant
VISUALIZATION_EXTENSION_POINTSThe renderer signatures FJ9 must be built around so that adding a partially observable layer later does not require rewriting it. Recorded here, in FJ10, because that is the whole reason this gate runs before the visualisation one.
A renderer whose only entry point takes a DuckieWorldState hardens the assumption that the thing being drawn is the latent truth. Belief-space visualisation is precisely the case where that is false.
Duckietown.ObservabilityClass — Type
ObservabilityClassHow a component of the 15-D privileged feature vector could ever be obtained:
SENSOR_ESTIMABLE— a camera or encoder could estimate it, with error.TEMPORALLY_DERIVED— needs tracking across frames, not one observation.MAP_PRIVILEGED— needs ground-truth map geometry beyond sensing range.SIMULATOR_PRIVILEGED— simulator bookkeeping, unobservable in principle.AGENT_MEMORY— the agent's own internal memory; belongs in the belief or the agent state, and is not an observation at all.
Duckietown.ReadinessItem — Type
ReadinessItemOne audited component: its status, the evidence that produced it (a probe of the package, not an opinion) and the change required.
Duckietown.ReadinessStatus — Type
ReadinessStatusREADY — usable as-is by a partially observable formulation. NEEDS_REFACTOR — present but must change first; the change is named. NOT_READY — absent. What has to be built is named.
Duckietown._has_duckietown_method — Method
Does mod.name carry any method mentioning one of this package's types? Guarded so a probe can never make the audit throw.
Duckietown.continuous_state_observability — Method
continuous_state_observability() -> Vector{ComponentObservability}Classify every component of ContinuousState. This is the concrete evidence that the 15-D vector is a privileged feature projection and not an observation: only a minority of its components could come from a sensor at all, and two of them are the agent's own memory.
Duckietown.observability_counts — Function
observability_counts(rows=continuous_state_observability()) -> NamedTupleDuckietown.observability_table — Function
observability_table(rows=continuous_state_observability()) -> StringDuckietown.pomdp_readiness — Function
pomdp_readiness(mdp) -> Vector{ReadinessItem}Probe the package for everything a partially observable formulation needs.
Determined by inspection of live types and method tables, so the result tracks the code rather than the documentation.
Duckietown.readiness_counts — Method
readiness_counts(items) -> NamedTupleDuckietown.readiness_table — Method
readiness_table(items) -> StringDuckietown.DuckieActionSpace — Type
DuckieActionSpaceContinuous action box [0, v_fast] x [-w0, w0] (m/s, rad/s), matching the reference ContinuousDuckieMDPEnv.action_space. rand(rng, space) samples uniformly — this is what a progressive-widening planner draws from.
Duckietown.DuckieInitialStateDistribution — Type
DuckieInitialStateDistributionThe initial-state distribution rho_0 (DuckieMDPEnv.reset). Sampling is implicit: rand(rng, d) runs the reference spawn loop — up to spawn_attempts curriculum attempts, each sampling a pose on the start tile (Simulator.reset), rebuilding the world, and testing the wrapper's acceptance predicate (|d|, |phi|, position bounds, route direction). The last candidate is returned if none is accepted, matching the reference RuntimeError case being unreachable in the shipped configs; pass strict = true to raise instead.
Duckietown.DuckietownMDP — Type
DuckietownMDP{A} <: MDP{DuckieWorldState, A}The Duckietown driving task as a POMDPs.jl MDP over the canonical branchable world state.
A = MacroAction: discrete 7-action problem (Q-learning/SARSA/MCTS).A = DuckieAction: continuous[v_cmd, omega_cmd]problem (SAC/TD3/DPW).
Construct from an experiment YAML (the reference config is the single source of every parameter):
mdp = DuckietownMDP("../duckduck/policies/q_learning/training_config.yaml")
mdpc = DuckietownMDP("../duckduck/policies/sac/training_config.yaml";
action_space = :continuous)
s0 = rand(rng, initialstate(mdp))
x = gen(mdp, s0, FAST_STRAIGHT, rng) # x.sp, x.rsimulate_decision(mdp.transition, s, a, rng) remains available for the full TransitionResult (reward breakdown, events, reason, projections) — gen deliberately exposes only (sp, r).
Duckietown.DuckietownMDP — Method
DuckietownMDP(config; action_space=:discrete, map=initial_map(config),
discount=config.solver.gamma)Build the MDP from a loaded DuckietownConfig. The discrete variant restricts the action set to the solver's allowed_actions when the config declares them (tabular experiments use 0:6, i.e. all seven).
Duckietown.build_world — Method
build_world(mdp, pos, angle) -> DuckieWorldStateA fresh world at the given ego pose: ego at rest with an empty command window, the injected duckie in its reset condition, the map's stop signs, and cleared stop/lane memory (DuckieMDPEnv.reset sets _mdp_sigma_stop = false and _mdp_last_lane_position = (1.0, 1.0), and stop_tracker.reset()).
Duckietown.is_truncated — Method
is_truncated(mdp, s) -> BoolHorizon truncation (step_count >= max_steps) for the same state.
Duckietown.spawn_accepted — Method
spawn_accepted(mdp, world, raw) -> BoolDuckieMDPEnv._spawn_is_accepted: the curriculum limits on |d| and |phi|, the optional x-z spawn rectangle, and the optional route-direction alignment.
POMDPs.gen — Method
POMDPs.gen(mdp, s, a, rng) -> (sp = ..., r = ...)Thin adapter over simulate_decision: one macro-decision (frame_skip physics ticks under the locked transition order), with the stochastic pedestrian trigger drawn from rng. s is never mutated, so the same state may be branched with different actions.
POMDPs.isterminal — Method
POMDPs.isterminal(mdp, s) -> Booltrue only for a GENUINE terminal (duck_collision, other_collision, offroad, goal) — the cases that break TD bootstrapping. A timeout is truncation imposed by the experiment horizon, not an absorbing physical state, so it is deliberately NOT terminal here; use is_truncated (or termination_reason) for the horizon, exactly as the reference wrapper separates terminated from truncated.
Duckietown.KNOWN_LIMITATIONS — Constant
KNOWN_LIMITATIONSWhat is deliberately not done, recorded so the manifest cannot imply otherwise. Omitting a deferred decision from a reproducibility statement is the same class of error as a stale claim.
Duckietown.SOURCE_IMPORT_BAN — Constant
SOURCE_IMPORT_BANPackages src/ must never import. The core is usable with none of them installed; each is a weak dependency served by an extension.
Duckietown.STALE_CLAIMS — Constant
STALE_CLAIMSClaims the evidence has contradicted. Each is banned from the normative documents; the allowlist names the files permitted to quote it, which are the correction itself and the tests that guard it.
FJ9.6 is the reason this exists: docs/src/validation/FJ8_STATUS.md asserted that TD3 "never reaches a stop sign" while its own artefact recorded 2 289 stop-zone decisions. The sentence survived because nothing checked prose.
Duckietown.ArtifactRecord — Type
ArtifactRecordDuckietown.ArtifactStatus — Type
ArtifactStatusHow an artefact comes to exist, which determines what "reproduce" means for it.
REBUILT — regenerated from source data on every run; a figure or a report. PERSISTED_SOURCE — the recorded evidence itself. Re-running the experiment that produced it is a different experiment, so this is checked, never rebuilt. PROVISIONED_FROZEN_INPUT — extracted once from a read-only upstream checkpoint by a step deliberately kept off the main path.
Duckietown.DocIssue — Type
DocIssueDuckietown.artifact_ledger — Method
artifact_ledger(root) -> Vector{ArtifactRecord}Every artefact, what kind of thing it is, and whether it is there (the FJ9.9c ledger).
The distinction matters: a PERSISTED_SOURCE that a rebuild would overwrite is not reproducibility, it is data loss. Only REBUILT entries are expected to be regenerable.
Duckietown.core_fingerprint — Method
core_fingerprint(mdp) -> StringAn identity for the FORMULATION (FJ9.9e): action semantics, state semantics, reward configuration and discount.
Loading Makie, MCTS or PythonCall must not change it. FJ8.5 established that solver integrations live in extensions; this makes the claim measurable — compute it with each optional package loaded and compare.
Duckietown.documentation_audit — Method
documentation_audit(root) -> Vector{DocIssue}Executable documentation consistency (FJ9.9d).
Checks three things across every .md and .jl in the repository:
- no
STALE_CLAIMSoutside their allowlist; - every markdown link to a repository path resolves;
- every backticked
artifacts/...ordocs/...path that looks like a file actually exists.
Duckietown.source_import_audit — Method
source_import_audit(root) -> Vector{DocIssue}Lint src/ for a banned import (FJ9.9e). Importing a planning library in the core would make the package refuse to load without that library installed, which is the architecture FJ8 was rebuilt to avoid.
The banned tokens are assembled at runtime rather than written out, because FJ8.1 and FJ8.5 lint src/ for solver vocabulary and a doc comment spelling one out is indistinguishable, to them, from the real thing. Those guards are stricter than this one and have no allowlist; that is the right trade.