Interfaces

POMDPs.jl, the RL-style environment, planning contract, reproducibility.

Duckietown.MDPLike — Type
MDPLike{A}

The model or a transparent wrapper around it, for a given action type. Code that only reads the model accepts this so a diagnostic wrapper never forces a second implementation of anything.

source
Duckietown.InstrumentedMDP — Type
InstrumentedMDP(mdp)

The same model, counting generative calls. It forwards every POMDPs.jl function and — through getproperty — every field, so anything written against DuckietownMDP works on it unchanged.

It is a measuring device, not a translation layer: it must not, and does not, alter states, rewards, transitions, action bounds or termination. That transparency is asserted bitwise in test/test_fj8_planning.jl rather than merely claimed here.

Only calls through POMDPs.gen are counted, which is exactly a planner's consumption of the model — the evaluator's own environment step goes through simulate_decision and is deliberately not charged as planning cost.

source
Duckietown.PlanningDiagnostics — Type
PlanningDiagnostics(planning_time, model_calls, extra)

What producing one action cost. Only two fields are universal:

  • planning_time — wall-clock seconds for the decision;
  • model_calls — generative-model calls consumed, or -1 when not measured.

Everything solver-specific goes in extra, an open NamedTuple. A tree search may report (tree_nodes = ..., max_depth = ..., iterations = ...); a value-iteration style solver (backups = ..., convergence_iterations = ...); a particle method (particles = ..., belief_nodes = ...). The evaluator never looks inside, so it cannot assume a planner has a tree — or has any internal structure at all, which is the normal case for a learned policy.

source
Duckietown._probe_states — Method

States along one driven trajectory, with the action taken at each, used so capability probing is not confined to the spawn state. Returns (states, actions).

source
Duckietown.model_calls — Method
model_calls(mdp) -> Int

Generative calls consumed so far; -1 for an uninstrumented model, so callers report "not measured" rather than a misleading zero.

source
Duckietown.model_capabilities — Method
model_capabilities(mdp; seeds) -> NamedTuple

What this model actually offers, determined by exercising the interface rather than by asserting it. The question a new solver should raise is "are this solver's requirements met by the model?", not "can the model be bent to fit the solver?" — this is the answer to the first form.

Stochasticity is reported as three separate facts, because this model is conditionally stochastic and collapsing that into one flag would mislead a planner:

  • consumes_rng — the transition draws from the caller's stream at some reachable state;
  • stochastic_state_fraction — at what fraction of the probed states it does so. The pedestrian trigger only fires when a duck is armed, ahead and inside the trigger distance window, so most states have a deterministic successor;
  • stochastic_outcomes — whether two seeds were actually observed to produce different successors.

A planner that widens over states (DPW's k_state/alpha_state) is doing useful work only in the fraction of states counted here; elsewhere the successor is a function of (s, a) alone. A false is a statement about the states and seeds probed, never a proof of determinism.

Pass policy to drive the probe trajectory with a competent controller. This matters: under a constant action the vehicle leaves the road within a few decisions and never reaches the states where a pedestrian can trigger, so the default probe under-reports what the model can do. The policy is only a way of reaching representative states — no property reported here depends on which policy is used, only on which states it visits.

source
Duckietown.plan_action — Method
plan_action(policy, mdp, s) -> (action, PlanningDiagnostics)

One decision plus what it cost. The default times policy_action and reads the model's call counter, which already works for any solver without the solver knowing this package exists.

A solver extension may add a method that fills richer extra fields. That is the only thing an extension is ever expected to add, and it is optional.

source
Duckietown.policy_action — Method
policy_action(policy, mdp, s) -> action

Ask any policy for its action, reconciling the two calling conventions in play — inside this package, without adding a method to POMDPs.action.

POMDPs.jl's contract is action(policy, x): solve has already bound the model into the policy, so a planner needs no model argument. This package's evaluator passes the model explicitly, because a stateless policy (a Q-table, an actor network) is reusable across models and needs it.

Extending POMDPs.action with a three-argument form would change how a generic POMDPs.jl policy behaves for everyone who loads this package — the opposite of the goal, which is that solvers plug in without the model altering the ecosystem around them. So the adaptation lives here, on a function this package owns:

  • a POMDPs.Policy (anything solve returns) is asked the standard way;
  • an AbstractPolicy (this package's tabular and actor adapters) is asked with the model, which is how those are defined;
  • anything else falls back to the three-argument form, so a hand-written policy that defines it keeps working.
source
Duckietown.AbstractPolicy — Type
AbstractPolicy

Interface boundary for reference-policy adapters (solvers/adapters.jl) and future solver policies. An adapter wraps one shipped artefact (policies/*/policy.npy or policy.pt) and maps the appropriate projection (discrete or encoded continuous observation) to an action.

source
Duckietown.act — Function
act(policy, observation, rng) -> action

Return the policy's action for the given observation (post-encoding continuous vector for SAC/TD3 adapters, raw state for tabular adapters). rng is required for stochastic exploration behaviour; deterministic evaluation must not consume it.

source
Duckietown.VISUALIZATION_EXTENSION_POINTS — Constant
VISUALIZATION_EXTENSION_POINTS

The renderer signatures FJ9 must be built around so that adding a partially observable layer later does not require rewriting it. Recorded here, in FJ10, because that is the whole reason this gate runs before the visualisation one.

A renderer whose only entry point takes a DuckieWorldState hardens the assumption that the thing being drawn is the latent truth. Belief-space visualisation is precisely the case where that is false.

source
Duckietown.ObservabilityClass — Type
ObservabilityClass

How a component of the 15-D privileged feature vector could ever be obtained:

  • SENSOR_ESTIMABLE — a camera or encoder could estimate it, with error.
  • TEMPORALLY_DERIVED — needs tracking across frames, not one observation.
  • MAP_PRIVILEGED — needs ground-truth map geometry beyond sensing range.
  • SIMULATOR_PRIVILEGED — simulator bookkeeping, unobservable in principle.
  • AGENT_MEMORY — the agent's own internal memory; belongs in the belief or the agent state, and is not an observation at all.
source
Duckietown.ReadinessItem — Type
ReadinessItem

One audited component: its status, the evidence that produced it (a probe of the package, not an opinion) and the change required.

source
Duckietown.ReadinessStatus — Type
ReadinessStatus

READY — usable as-is by a partially observable formulation. NEEDS_REFACTOR — present but must change first; the change is named. NOT_READY — absent. What has to be built is named.

source
Duckietown.continuous_state_observability — Method
continuous_state_observability() -> Vector{ComponentObservability}

Classify every component of ContinuousState. This is the concrete evidence that the 15-D vector is a privileged feature projection and not an observation: only a minority of its components could come from a sensor at all, and two of them are the agent's own memory.

source
Duckietown.pomdp_readiness — Function
pomdp_readiness(mdp) -> Vector{ReadinessItem}

Probe the package for everything a partially observable formulation needs.

Determined by inspection of live types and method tables, so the result tracks the code rather than the documentation.

source
Duckietown.DuckieActionSpace — Type
DuckieActionSpace

Continuous action box [0, v_fast] x [-w0, w0] (m/s, rad/s), matching the reference ContinuousDuckieMDPEnv.action_space. rand(rng, space) samples uniformly — this is what a progressive-widening planner draws from.

source
Duckietown.DuckieInitialStateDistribution — Type
DuckieInitialStateDistribution

The initial-state distribution rho_0 (DuckieMDPEnv.reset). Sampling is implicit: rand(rng, d) runs the reference spawn loop — up to spawn_attempts curriculum attempts, each sampling a pose on the start tile (Simulator.reset), rebuilding the world, and testing the wrapper's acceptance predicate (|d|, |phi|, position bounds, route direction). The last candidate is returned if none is accepted, matching the reference RuntimeError case being unreachable in the shipped configs; pass strict = true to raise instead.

source
Duckietown.DuckietownMDP — Type
DuckietownMDP{A} <: MDP{DuckieWorldState, A}

The Duckietown driving task as a POMDPs.jl MDP over the canonical branchable world state.

  • A = MacroAction: discrete 7-action problem (Q-learning/SARSA/MCTS).
  • A = DuckieAction: continuous [v_cmd, omega_cmd] problem (SAC/TD3/DPW).

Construct from an experiment YAML (the reference config is the single source of every parameter):

mdp  = DuckietownMDP("../duckduck/policies/q_learning/training_config.yaml")
mdpc = DuckietownMDP("../duckduck/policies/sac/training_config.yaml";
                     action_space = :continuous)

s0 = rand(rng, initialstate(mdp))
x  = gen(mdp, s0, FAST_STRAIGHT, rng)   # x.sp, x.r

simulate_decision(mdp.transition, s, a, rng) remains available for the full TransitionResult (reward breakdown, events, reason, projections) — gen deliberately exposes only (sp, r).

source
Duckietown.DuckietownMDP — Method
DuckietownMDP(config; action_space=:discrete, map=initial_map(config),
              discount=config.solver.gamma)

Build the MDP from a loaded DuckietownConfig. The discrete variant restricts the action set to the solver's allowed_actions when the config declares them (tabular experiments use 0:6, i.e. all seven).

source
Duckietown.build_world — Method
build_world(mdp, pos, angle) -> DuckieWorldState

A fresh world at the given ego pose: ego at rest with an empty command window, the injected duckie in its reset condition, the map's stop signs, and cleared stop/lane memory (DuckieMDPEnv.reset sets _mdp_sigma_stop = false and _mdp_last_lane_position = (1.0, 1.0), and stop_tracker.reset()).

source
Duckietown.spawn_accepted — Method
spawn_accepted(mdp, world, raw) -> Bool

DuckieMDPEnv._spawn_is_accepted: the curriculum limits on |d| and |phi|, the optional x-z spawn rectangle, and the optional route-direction alignment.

source
POMDPs.gen — Method
POMDPs.gen(mdp, s, a, rng) -> (sp = ..., r = ...)

Thin adapter over simulate_decision: one macro-decision (frame_skip physics ticks under the locked transition order), with the stochastic pedestrian trigger drawn from rng. s is never mutated, so the same state may be branched with different actions.

source
POMDPs.isterminal — Method
POMDPs.isterminal(mdp, s) -> Bool

true only for a GENUINE terminal (duck_collision, other_collision, offroad, goal) — the cases that break TD bootstrapping. A timeout is truncation imposed by the experiment horizon, not an absorbing physical state, so it is deliberately NOT terminal here; use is_truncated (or termination_reason) for the horizon, exactly as the reference wrapper separates terminated from truncated.

source
Duckietown.KNOWN_LIMITATIONS — Constant
KNOWN_LIMITATIONS

What is deliberately not done, recorded so the manifest cannot imply otherwise. Omitting a deferred decision from a reproducibility statement is the same class of error as a stale claim.

source
Duckietown.SOURCE_IMPORT_BAN — Constant
SOURCE_IMPORT_BAN

Packages src/ must never import. The core is usable with none of them installed; each is a weak dependency served by an extension.

source
Duckietown.STALE_CLAIMS — Constant
STALE_CLAIMS

Claims the evidence has contradicted. Each is banned from the normative documents; the allowlist names the files permitted to quote it, which are the correction itself and the tests that guard it.

FJ9.6 is the reason this exists: docs/src/validation/FJ8_STATUS.md asserted that TD3 "never reaches a stop sign" while its own artefact recorded 2 289 stop-zone decisions. The sentence survived because nothing checked prose.

source
Duckietown.ArtifactStatus — Type
ArtifactStatus

How an artefact comes to exist, which determines what "reproduce" means for it.

REBUILT — regenerated from source data on every run; a figure or a report. PERSISTED_SOURCE — the recorded evidence itself. Re-running the experiment that produced it is a different experiment, so this is checked, never rebuilt. PROVISIONED_FROZEN_INPUT — extracted once from a read-only upstream checkpoint by a step deliberately kept off the main path.

source
Duckietown.artifact_ledger — Method
artifact_ledger(root) -> Vector{ArtifactRecord}

Every artefact, what kind of thing it is, and whether it is there (the FJ9.9c ledger).

The distinction matters: a PERSISTED_SOURCE that a rebuild would overwrite is not reproducibility, it is data loss. Only REBUILT entries are expected to be regenerable.

source
Duckietown.core_fingerprint — Method
core_fingerprint(mdp) -> String

An identity for the FORMULATION (FJ9.9e): action semantics, state semantics, reward configuration and discount.

Loading Makie, MCTS or PythonCall must not change it. FJ8.5 established that solver integrations live in extensions; this makes the claim measurable — compute it with each optional package loaded and compare.

source
Duckietown.documentation_audit — Method
documentation_audit(root) -> Vector{DocIssue}

Executable documentation consistency (FJ9.9d).

Checks three things across every .md and .jl in the repository:

  • no STALE_CLAIMS outside their allowlist;
  • every markdown link to a repository path resolves;
  • every backticked artifacts/... or docs/... path that looks like a file actually exists.
source
Duckietown.source_import_audit — Method
source_import_audit(root) -> Vector{DocIssue}

Lint src/ for a banned import (FJ9.9e). Importing a planning library in the core would make the package refuse to load without that library installed, which is the architecture FJ8 was rebuilt to avoid.

The banned tokens are assembled at runtime rather than written out, because FJ8.1 and FJ8.5 lint src/ for solver vocabulary and a doc comment spelling one out is indistinguishable, to them, from the real thing. Those guards are stricter than this one and have no allowlist; that is the right trade.

source