Frozen policies
Native inference for the reference tabular and actor policies.
Duckietown.LinearLayer — Type
LinearLayerOne torch.nn.Linear: y = W * x + b with W of size (out, in).
Duckietown.SACActorPolicy — Type
SACActorPolicy <: AbstractPolicyNative SAC actor (deterministic evaluation).
Duckietown.SACActorPolicy — Method
SACActorPolicy(dir)Load from a directory of exported parameter files (backbone_0_weight.npy, backbone_0_bias.npy, backbone_2_*, mean_*, log_std_*, action_scale.npy, action_bias.npy).
Duckietown.TD3ActorPolicy — Type
TD3ActorPolicy <: AbstractPolicyNative TD3 actor (deterministic by construction).
Duckietown.TD3ActorPolicy — Method
TD3ActorPolicy(dir; action_low, action_high)Duckietown.act — Function
act(policy, obs[, rng]) -> DuckieActionThe evaluation action: for SAC tanh(mean)*scale+bias; for TD3 the same followed by the agent's clip. Deterministic — rng is never consumed.
Duckietown.forward — Method
forward(policy, obs) -> NamedTupleFull activation trace. Field names match the oracle's response keys so the two can be compared layer by layer.
POMDPs.action — Method
POMDPs.action(policy, mdp, s) -> DuckieActionDrive a learned continuous policy from a world state: project to the 15-component continuous state and encode it exactly as the reference does, then run the actor.
Duckietown.TIE_ATOL — Constant
TIE_ATOLAbsolute tolerance of the reference's near-tie window (np.isclose(..., rtol=0.0, atol=1e-12)).
Duckietown.QDecision — Type
QDecisionOne greedy decision plus the evidence the reference records with it: the full Q-value row, the near-tied action set, and the top-1/top-2 margin.
Duckietown.QTablePolicy — Type
QTablePolicy <: AbstractPolicyA shipped tabular reference policy (Q-learning or SARSA) exposed through the validated MDP. Holds the Q-table exactly as stored (Float32, shape Q_SHAPE), the allowed action ids, and the action table the ids map to.
pol = QTablePolicy("../duckduck/policies/q_learning/policy.npy")
a = act(pol, raw_state) # MacroAction
d = decide(pol, discretize(raw)) # full decision record incl. ties/marginDuckietown.QTablePolicy — Method
QTablePolicy(path; allowed_actions=0:6, action_cfg=ActionConfig(), solver=:q_learning)Load a policy.npy checkpoint. Validates the shape against Q_SHAPE, that every entry is finite, and that the allowed action ids are unique and inside 0:6 — the same guards the reference adapter applies.
Duckietown.act — Function
act(policy, raw_state[, rng]) -> MacroActionDeterministic greedy action. The rng is accepted for interface uniformity and deliberately never consumed.
Duckietown.all_state_indices — Method
all_state_indices() -> Vector{NTuple{7,Int}}Every valid 0-based discrete state index, in the reference's C-order enumeration (last dimension varies fastest) — 9 000 of them.
Duckietown.decide — Method
decide(policy, index) -> QDecisionGreedy decision for a 0-based discrete state index (as produced by discretize), following the reference rule verbatim.
Duckietown.greedy_action_table — Method
greedy_action_table(policy) -> Vector{Int}The selected action id for every discrete state, in all_state_indices order. This is the object FJ7 compares against the reference.
Duckietown.read_npy — Method
read_npy(path) -> ArrayMinimal reader for the NumPy .npy v1/v2 format restricted to what the shipped checkpoints use: plain (non-pickled) little-endian numeric arrays. Returns a Julia array with the file's own shape, honouring fortran_order.
Duckietown.tie_statistics — Method
tie_statistics(policy) -> NamedTupleHow often the near-tie window actually decides something: the number of states with more than one tied action, how often the tie changes the outcome relative to a plain argmax, and the margin distribution.
POMDPs.action — Method
POMDPs.action(policy, mdp, s) -> MacroActionDrive the policy directly from a world state: project with the validated observer, discretize, then decide.