Backends

Native and reference backends.

Duckietown.AbstractBackend — Type
AbstractBackend

Interface boundary between the decision-process formulation and the two simulation backends (FJ0 audit section E/F):

  • native_julia.jl — full generative reproduction (FJ3 dynamics);
  • gym_duckietown.jl — PythonCall bridge to the unmodified duckduck DuckieMDPEnv/ContinuousDuckieMDPEnv (reference, for parity).

Backends translate between DuckieWorldState (canonical branchable dynamics state) and their internal representation, and expose the decision step with the locked transition order:

1. duck controller `before_step` (activation draw)
2. action → wheel commands
3. `frame_skip` delayed-DB18 physics ticks (duckie step included)
4. raw-state extraction (lane frame, stop candidate, duck threat)
5. StopTracker update (sigma, events)
6. collision / termination classification (duck > other > timeout >
   offroad > goal > in_progress)
7. reward evaluation on the post-transition state

The projections get_raw_state/get_continuous_state are pure functions of the world state (user constraint #2: the 7-D/15-D states are projections, not the canonical MDP state).

source
Duckietown.reset! — Function
reset!(backend, seed) -> DuckieWorldState

Sample s0 from ρ0 (spawn curriculum, duck controller reset, stop memory reset) and return the canonical world state.

source
Duckietown.step! — Function
step!(backend, world::DuckieWorldState, action, rng) ->
    (world′, reward, terminated, truncated, info)

One decision: frame_skip physics ticks under the locked transition order. info carries events, reward breakdown, raw/continuous projections, and termination reason.

source
Duckietown.AbstractReferenceBackend — Type
AbstractReferenceBackend <: AbstractBackend

Common supertype of the reference (Python) backends. Two transports exist for the SAME reference runtime and the same Session semantics:

  • ProcessReferenceBackend — out-of-process JSON-lines server (tools/parity/reference_server.py). Works from a Windows Julia against a WSL Python; this is what FJ5 validated.
  • PythonCallReferenceBackend — in-process via PythonCall, available when Julia and the reference Python live on the same platform (FJ5-R). Provided by the DuckietownPythonCallExt package extension, so PythonCall stays an optional dependency: using Duckietown never touches Python.

Both expose the same verbs — ref_reset!, ref_get_state, ref_set_state!, ref_step!, ref_probe_stop, close — and share one state mapping (world_to_ref / ref_to_world), so callers only choose at construction.

source
Duckietown.ProcessReferenceBackend — Type
ProcessReferenceBackend <: AbstractReferenceBackend

Handle on a running reference-server process.

ref = ReferenceBackend("q_learning"; seed = 53)
s, _ = ref_reset!(ref, 53)                # DuckieWorldState from the reference
ref_set_state!(ref, world)                # inject a Julia state
sp, dump = ref_step!(ref, FAST_STRAIGHT)  # reference transition from it
close(ref)

reference_backend_available reports whether the backend can start at all (WSL + the ddm-ref env present), so parity test sets skip cleanly on machines without the reference environment.

source
Duckietown.ProcessReferenceBackend — Method
ReferenceBackend(config; seed, action_space=:discrete, overrides=Dict())

Launch the reference server and build the reference env from the experiment YAML config ("q_learning", "sarsa", "sac", "td3"). overrides is a nested section => Dict(key => value) mapping applied on top of the YAML — used only for explicitly-labelled variant scenarios; the baseline configs on disk are never modified.

source
Base.close — Method
close(backend)

Shut the reference process down cleanly: ask it to quit, close its stdin so the server's for line in sys.stdin loop ends, then reap it.

close(::Base.Process) is NOT sufficient here — it returns successfully while process_running stays true, leaving an orphaned child holding a live pipe. On this Windows build that half-closed state later crashed the GC (EXCEPTION_ACCESS_VIOLATION in gc_mark_stack) once a second backend was created in the same session. Closing proc.in explicitly and waiting is what actually terminates it.

source
Duckietown.PythonCallReferenceBackend — Method
PythonCallReferenceBackend(config; seed, action_space, overrides, map)

In-process reference backend (FJ5-R). Requires the PythonCall extension:

using PythonCall                      # loads DuckietownPythonCallExt
ref = PythonCallReferenceBackend("q_learning"; seed = 53)

Without PythonCall loaded this throws a message saying exactly that.

source
Duckietown.ref_call — Method
ref_call(backend, message) -> result

Send one protocol message and return its result, raising the reference traceback as an ErrorException on failure.

source
Duckietown.ref_probe_stop — Method
ref_probe_stop(backend; decisions=400, policy="lane_follow") -> rows

Record every stop-candidate filter quantity per decision from the LIVE reference runtime (observation only; the baseline config is untouched).

source
Duckietown.ref_step! — Method
ref_step!(backend, action) -> (world, dump)

One reference decision (DuckieMDPEnv.step / ContinuousDuckieMDPEnv.step). action is a MacroAction/Int for the discrete env or a DuckieAction/vector for the continuous one. dump.result carries the reward breakdown, events, termination reason and wheel commands.

source
Duckietown.ref_to_world — Method
ref_to_world(dump, map) -> DuckieWorldState

Import a reference state dump into the canonical Julia world state. Every dynamics-relevant field crosses: the DB18 pose/velocity matrices q0/v0, the wheel-axis angles, the trimmed delayed-command window, each duckie's full object state, the controller counters, and the stop/lane memory.

source
Duckietown.world_to_ref — Method
world_to_ref(world) -> Dict

Export a canonical world state in the reference server's schema, so the same latent state x_t can be injected into the real Python simulator (set_state) and stepped side by side with the native model.

source
Duckietown.TorchPolicyReferenceBackend — Type
TorchPolicyReferenceBackend

Handle on the running torch_policy_server.py process. It answers init (metadata + exports the actor weights as .npy) and infer (per-layer activations for one 15-D observation).

source
Duckietown.TorchPolicyReferenceBackend — Method
TorchPolicyReferenceBackend()

Launch the oracle. Uses the same clean-shutdown discipline as the simulator backend (close(proc.in) then reap — close(::Process) alone leaves the child running).

source
Duckietown.torch_policy_available — Method
torch_policy_available() -> Bool

Whether the ddm-torch oracle can be started (WSL reachable and the env present). Parity tests skip cleanly when it cannot.

source
Duckietown.torch_policy_infer — Method
torch_policy_infer(backend, policy, obs) -> Dict

Reference activations for one 15-D observation: every linear/ReLU output, the head outputs, the squashed action, the scaled action, the action the agent's own select_action(deterministic=true) returns, and the clipped action.

source
Duckietown.torch_policy_init! — Method
torch_policy_init!(backend, policy) -> Dict

Load "sac" or "td3" in the oracle and return its metadata: obs_dim, hidden, action bounds, parameter shapes, the exported weight files, the torch/numpy/python versions and the evaluation rule read from the reference agent source.

source