In LeRobot v0.6.1, chunk_size is how many actions the model predicts per call and n_action_steps is how many of them the robot plays before the policy looks again. The defaults set them equal (ACT 100/100, SmolVLA, pi0 and pi0.5 50/50), so each call's whole chunk runs open loop and the policy sees one observation per chunk. The papers behind these policies report executing less: the pi0 authors ran 16 or 25 of 50, Diffusion Policy found an action horizon of 8 best for most tasks, ACT recommends temporal ensembling, and SmolVLA's own ablation drops from 82.8% average LIBERO success at 10 executed steps to 51.8% at 50. LeRobot issue 4614 reports the shipped lerobot/smolvla_libero config (50 of 50) costing about 20 points against n_action_steps=10, and one contributor's sweep on libero_spatial went from 124 of 150 at 1 to 73 of 150 at 50. chunk_size is baked in at training; n_action_steps is a CLI override on top of config.json in both lerobot-eval and lerobot-rollout. On the async path it is bypassed entirely. Start by sweeping --policy.n_action_steps over 1, 10, 25 and the full chunk on one checkpoint with a fixed seed.
LeRobot n_action_steps and chunk_size look like one setting written twice, and the defaults encourage that reading: ACT ships 100 and 100, SmolVLA, pi0 and pi0.5 ship 50 and 50. They are two different decisions. chunk_size is how many future actions the model predicts per forward pass. n_action_steps is how many of those the robot actually plays before the policy looks at the world again. Set them equal and every chunk runs open loop to the end.
The papers these policies come from report other settings. The pi0 authors executed 16 of 50 actions at 20 Hz and 25 of 50 at 50 Hz. Diffusion Policy found an 8-step execution horizon best for most tasks. ACT's authors query every step and average overlapping predictions. SmolVLA's own ablation on LIBERO averaged 82.8% when executing 10 of a 50-step chunk and 51.8% when executing all 50.
Everything below is read from LeRobot v0.6.1, released 2026-08-03 and still the current release. The short version: chunk_size is a training decision you live with, n_action_steps is a deployment decision you should make on purpose, and the shipped default makes it for you.
chunk_size is what the model predicts, n_action_steps is what the robot plays
The ACT config docstring in v0.6.1 defines n_action_steps as the number of action steps to run in the environment for one invocation of the policy, no greater than the chunk size. Its worked example: with a chunk of 100 you may set this to 50, which means the model predicts 100 steps worth of actions, runs 50 in the environment, and throws the other 50 out. The pi0 config comments say the same in fewer words: chunk_size is the number of action steps to predict, which openpi calls action_horizon, and n_action_steps is the number to execute.
The two knobs live at different points in the model's life. For ACT, chunk_size is architectural: the decoder builds nn.Embedding(config.chunk_size, config.dim_model), so it sets a weight shape and cannot change on a trained checkpoint. For SmolVLA and pi0 it shapes the noise tensor and the attention mask rather than a parameter, but the released models were trained at 50, and running them at another value is untested. n_action_steps has no weights behind it at all. It is a runtime number stored in config.json.
Diffusion is the naming trap. LeRobot's Diffusion config has no chunk_size field; the prediction length is horizon. The Diffusion Policy paper uses three horizons: T_o for observations, T_p for prediction, and T_a, the action execution horizon, defined as the steps executed on the robot without re-planning. T_a is what LeRobot calls n_action_steps. When a Diffusion Policy discussion says "action horizon of 8", it means n_action_steps=8, not horizon=8.
Four of the five defaults execute the entire prediction. Diffusion is the exception, and at 32 of 64 it still executes half.
| Policy (v0.6.1) | Prediction length field | Default prediction | Default n_action_steps | n_obs_steps | Config-time check |
|---|---|---|---|---|---|
| ACT | chunk_size | 100 | 100 | 1 | ValueError if n_action_steps > chunk_size |
| SmolVLA | chunk_size | 50 | 50 | 1 | ValueError if n_action_steps > chunk_size |
| pi0 | chunk_size | 50 | 50 | 1 | ValueError if n_action_steps > chunk_size |
| pi0.5 | chunk_size | 50 | 50 | 1 | ValueError if n_action_steps > chunk_size |
| Diffusion | horizon | 64 | 32 | 2 | none on n_action_steps |
What select_action does with the action queue
In ACT, pi0 and pi0.5, reset() creates an action queue, a deque with maxlen=n_action_steps. On each call, select_action checks whether the queue is empty. If it is, it runs predict_action_chunk(batch)[:, :n_action_steps], transposes the result so time is the first axis, and extends the queue. Then it pops one action off the left and returns it. SmolVLA follows the same pattern and slices after the transpose. Diffusion adds observation queues that hold the last n_obs_steps observations, and its slice starts at index n_obs_steps - 1 rather than 0, because the first predicted steps line up with observations already seen.
The consequence is simple arithmetic. The policy sees a new observation only when the queue runs dry, which is once every n_action_steps ticks. At the defaults that is every 50 ticks for SmolVLA, pi0 and pi0.5, and every 100 ticks for ACT. LIBERO's environment runs at 20 Hz in LeRobot (the fps default, which the config notes must match robosuite's control frequency), so a SmolVLA policy at default settings replans every 2.5 seconds. At n_action_steps=10 it replans every 0.5 seconds with the same weights.
Both synchronous entry points go through this queue. lerobot-eval's rollout() calls policy.reset() and then policy.select_action(observation) every step. On hardware, lerobot-rollout's default backend is --inference.type=sync, documented as one policy call per control tick, and SyncInferenceEngine.get_action calls select_action. So the value you set governs replanning in simulation and on the robot alike, as long as you stay on the sync path.
The validators, and the combinations that raise
ACT and SmolVLA share one check in __post_init__, with this message when n_action_steps exceeds chunk_size:
ValueError: The chunk size is the upper bound for the number of action steps per model invocation. Got 60 for `n_action_steps` and 50 for `chunk_size`.
That is what --policy.n_action_steps=60 produces on a SmolVLA config with chunk_size 50. The comparison is strict greater-than, so equality (the shipped default) passes. pi0 and pi0.5 run the same comparison with different text: n_action_steps (60) cannot be greater than chunk_size (50).
ACT has a second check, and it is the one a single added flag runs into. When temporal_ensemble_coeff is set and n_action_steps is greater than 1, the config raises:
NotImplementedError: `n_action_steps` must be 1 when using temporal ensembling. This is because the policy needs to be queried every step to compute the ensembled action.
Any value of 2 or more fires it, including the default 100, and it runs before the chunk size check. Passing --policy.temporal_ensemble_coeff=0.01 alone on a stock ACT config therefore fails at startup. You need --policy.n_action_steps=1 in the same command. ACT also raises a ValueError if n_obs_steps is anything other than 1.
Diffusion has no check on n_action_steps at all. DiffusionPolicy.select_action's docstring states the requirement, n_action_steps <= horizon - n_obs_steps + 1, but __post_init__ checks the backbone, the noise scheduler, the image resize and crop settings, and that horizon divides evenly by the U-Net's downsampling factor, and nothing about n_action_steps. Our reading of the code: a value larger than 63 at the v0.6.1 defaults produces no error, and the slice that fills the queue quietly returns only the 63 actions that exist, so the policy replans sooner than the config says.
What the papers measured when they separated the two knobs
Four papers, four ways of executing less than the full prediction.
The caption's conclusion is that sampling new observations more frequently, every 1 or 10 steps, significantly improves performance, and the text adds that acting the entire chunk speeds up inference but reduces the robot's responsiveness to environmental changes. Table 12 ablates chunk size instead: 1 averaged 50.0, 10 averaged 84.0, 30 averaged 78.5, 50 averaged 80.3 and 100 averaged 74.5. Table 12's chunk-50 row and Table 13's one-step row match cell for cell, which fits the paper's implementation details: in simulation the authors sample a new observation and predict after every executed action. Our reading: Table 12 varies the chunk at one executed step, and Table 13 varies the executed steps at a chunk of 50. The same section states that the real-world evaluation used synchronous inference that executes the full chunk before sampling again, so the full-chunk setting is not absent from the paper; it is the one its own ablation scores lowest.
One caveat before you carry these numbers anywhere. Section 4.7 states that the ablations were trained from scratch without robotics pretraining, with the VLM backbone frozen and only the action expert trained. These are not the released checkpoint's scores. They show the direction of the effect, not its size on your model.
SmolVLA's ablation isolates the execution horizon
The SmolVLA paper ablates the execution horizon with the chunk held at 50, in Table 13, on LIBERO. Diffusion Policy also sweeps its execution horizon, in a figure rather than a table, covered below.
pi0, ACT and Diffusion Policy
The pi0 paper is explicit: the authors tried temporal ensembling early, found it hurt policy performance, and chose to execute chunks open loop, running inference every 0.8 seconds after 16 actions on the 20 Hz UR5e and Franka arms and every 0.5 seconds after 25 actions on the 50 Hz robots. With a 50-step chunk, that is never the whole chunk. They also report 73 ms of on-board inference on an RTX 4090, which is what makes replanning every 16 steps affordable. The ACT paper reports that chunking itself matters enormously: success went from 1% at k=1 to 44% at k=100 with temporal ensembling disabled. That ablation predicted and executed k steps together, so it measures chunking, not the execution horizon on its own. The separate temporal ensembling result is what bears on execution: a 3.3% gain for ACT. The Diffusion Policy paper states the trade-off in one line: too long an action horizon reduces performance through slow reaction time, and 8 steps was optimal for most tasks it tested. It also reports that the policy maintained peak performance with latency up to 4 steps. In the settings these papers chose or found best, each executes less than the full prediction. LeRobot's defaults for ACT, SmolVLA, pi0 and pi0.5 execute all of it.
| Paper | Prediction length | Executed per policy call | Reported finding |
|---|---|---|---|
| ACT (arXiv 2304.13705) | chunk of 100 | 1, with temporal ensembling | Temporal ensembling added 3.3% for ACT |
| pi0 (arXiv 2410.24164) | H = 50 | 16 at 20 Hz, 25 at 50 Hz | Temporal ensembling hurt; partial chunks run open loop |
| Diffusion Policy (arXiv 2303.04137) | T_p | T_a = 8 | 8 steps optimal for most tasks tested |
| SmolVLA (arXiv 2506.01844) | chunk of 50 | ablated 1, 10, 30, 50 | 10 steps averaged 82.8%, 50 steps 51.8% |
| Executed steps | Spatial | Object | Goal | Long (10) | Average |
|---|---|---|---|---|---|
| 1 | 89 | 94 | 85 | 53 | 80.3 |
| 10 | 89 | 94 | 91 | 57 | 82.8 |
| 30 | 76 | 91 | 74 | 42 | 70.8 |
| 50 | 54 | 70 | 58 | 25 | 51.8 |
The worked example: a shipped config that plays the whole chunk
The config that travels with a checkpoint matters more than the dataclass default, because --policy.path loads config.json and inherits whatever is stored there. Here is what several public checkpoints on the Hugging Face Hub stored as shipped in September 2026. These are mutable files, so check the one you pull.
Two SmolVLA LIBERO checkpoints under two organisations, one replanning every step and one every 50. The LeRobot LIBERO docs evaluate pi0.5 with --policy.n_action_steps=10, stating that this matches the original OpenPI implementation (the docs' claim; we did not check openpi), while the pi05_libero_finetuned config stores 50.
LeRobot issue 4614, still open, quantifies what the SmolVLA default costs. The reporter ran paired evaluations on lerobot 0.6.1 with only n_action_steps changed: on libero_object, 42 of 50 at 10 against 32 of 50 at 50 (McNemar p 0.031); on libero_spatial, 115 of 150 against 80 of 150 (p 2e-7). Forcing the HuggingFaceVLA checkpoint up to 50 moved it from 40 to 12 of 50 on object and 31 to 17 of 50 on spatial. A second contributor confirmed the direction on CUDA: 115 of 150 at 10 against 90 of 150 at 50 on libero_spatial, exact McNemar p 2.2e-05. A third ran a sweep on lerobot 0.6.1:
These are the contributors' runs, not ours. The shape is close to the SmolVLA paper's Table 13: modest differences from 1 to 10, a decline at 25 and a collapse at the full chunk. One difference: in this sweep 1 beat 10 by 10 episodes (McNemar p 0.013), where the paper's ablation had 10 slightly ahead.
The fix is one flag. lerobot-eval applies --policy.* overrides on top of the checkpoint's config.json, so the command below keeps the shipped weights and replaces the inherited 50 with 10. The --rename_map line maps this checkpoint's camera keys onto LIBERO's, and is needed for the command to run in 0.6.1.
lerobot-eval \
--policy.path=lerobot/smolvla_libero \
--policy.n_action_steps=10 \
--env.type=libero \
--env.task=libero_object \
--eval.n_episodes=50 \
--eval.batch_size=1 \
--seed=1000 \
--rename_map='{"observation.images.image": "observation.images.camera1", "observation.images.image2": "observation.images.camera2"}'Two things to know around this flag. First, loading with --policy.pretrained_path instead of --policy.path takes weights only and resets every stored setting to the dataclass default, which is why the LeRobot pi0.5 docs pass --policy.n_action_steps=10 explicitly; the full difference between the two flags is covered in our SmolVLA 0% success rate triage. Second, nothing warns you. The proposed fix, PR 4615, is open and unmerged, and it adds a docs recipe and a warning on LIBERO fps other than 20, not a warning on the horizon. The issue thread argued against a load-time warning on n_action_steps == chunk_size because that equality is the shipped default for ACT, pi0 and SmolVLA.
| Checkpoint | Prediction length | n_action_steps stored |
|---|---|---|
| lerobot/smolvla_libero | 50 | 50 |
| HuggingFaceVLA/smolvla_libero | 50 | 1 |
| lerobot/pi05_libero_base | 50 | 10 |
| lerobot/pi05_libero_finetuned | 50 | 50 |
| lerobot/diffusion_pusht | 16 (horizon) | 8 |
| lerobot/act_aloha_sim_transfer_cube_human | 100 | 100 |
n_action_steps | libero_spatial successes (10 tasks x 15 episodes) | Success rate |
|---|---|---|
| 1 | 124 / 150 | 82.7% |
| 5 | 115 / 150 | 76.7% |
| 10 | 114 / 150 | 76.0% |
| 25 | 106 / 150 | 70.7% |
| 50 | 73 / 150 | 48.7% |
ACT: temporal ensembling or a shorter horizon
ACT is the one LeRobot policy with a second way to replan often. Temporal ensembling, Algorithm 2 in the ACT paper, queries the policy at every control tick. Each call predicts a full chunk, so at any tick several past chunks each hold a prediction for the current timestep. The ensembler averages them with weights w_i = exp(-m * i), where w_0 is the oldest prediction. The paper states that a smaller m means faster incorporation of new observations.
In LeRobot v0.6.1 this is ACTTemporalEnsembler, switched on by temporal_ensemble_coeff. With it set, select_action bypasses the queue entirely and calls predict_action_chunk on every call. The ensembler's docstring spells out the sign convention: 0 weighs all predictions uniformly, positive values favour older ones, negative values favour newer ones. It also states that 0.01 is the value used by the original ACT work. That figure is LeRobot's attribution; the paper text gives the formula, not the number.
The cost is compute, not training. The ACT paper notes temporal ensembling incurs no additional training cost, only extra inference-time computation, and the extra is exactly one forward pass per control tick instead of one per n_action_steps ticks. At the default of 100, that is 100 times as many forward passes.
The temporal_ensemble_coeff field exists only in the ACT policy in v0.6.1. There is no temporal ensembling for SmolVLA, pi0, pi0.5 or Diffusion in LeRobot, and the pi0 authors found it hurt their policy anyway. For those, a shorter n_action_steps is the only lever.
On hardware, enabling it takes both flags. The command below follows the base-strategy example in the v0.6.1 lerobot-rollout docstring; add --robot.cameras with keys that match your policy's image features.
lerobot-rollout \ --strategy.type=base \ --policy.path=<your ACT checkpoint> \ --policy.temporal_ensemble_coeff=0.01 \ --policy.n_action_steps=1 \ --robot.type=koch_follower \ --robot.port=/dev/ttyACM0 \ --task="pick up cube" \ --duration=30
Drop the --policy.n_action_steps=1 line and the NotImplementedError above fires before the robot moves. For ACT, run both candidates through the same evaluation: ensembling at 0.01 with n_action_steps=1, and a plain queue at a shorter n_action_steps such as 10 or 25. The ACT authors' 3.3% is their policy on their tasks.
Picking n_action_steps on your own rig
for N in 1 10 25 50; do
lerobot-eval \
--policy.path=<your checkpoint> \
--policy.n_action_steps=$N \
--env.type=libero \
--env.task=libero_spatial \
--eval.n_episodes=50 \
--eval.batch_size=1 \
--seed=1000 \
--output_dir=./eval_logs/n_action_steps_$N
doneEvery row is a starting point for the sweep, not a replacement for it.
Sweep one knob with everything else fixed
Same checkpoint, same seed, same environment, one output directory per value. Values above the checkpoint's prediction length raise the ValueError, so the loop stops at the chunk size: 50 for SmolVLA and pi0, 100 for ACT. Add --rename_map if your checkpoint needs it. Treating the four runs as paired assumes each one sees the same initial states in the same order; confirm that on your setup before you run a paired test. How many trials a robot policy evaluation needs covers the McNemar comparison and the trial count required to separate two values honestly. The gaps reported in issue 4614 between 10 and 50 cleared significance at 50 and 150 episodes; the sweep's 5 against 10 (115 against 114 of 150, McNemar p 1.0) shows no detectable difference.
Check the latency bound before you go low
A smaller n_action_steps means more policy calls per second. On the sync path the tick that finds the queue empty also runs the forward pass, so if one call takes longer than the control period, that tick overruns and the robot pauses at every chunk boundary. At 20 Hz a control tick is 50 ms; pi0's reported 73 ms of on-board inference does not fit in one. Measure end-to-end latency on the deployment machine, as described in the sim-to-real gap post, before you settle on 1.
Account for the action representation and the async path
Delta and absolute action anchors change how errors accumulate inside an executed chunk, which moves the best execution horizon; the coupling is covered in joint versus end-effector action spaces. If you deploy with async inference, n_action_steps stops applying. The v0.6.1 policy server calls predict_action_chunk directly and truncates to actions_per_chunk, so select_action and its queue never run, and the replanning cadence is set by the client's chunk_size_threshold. Our VLA comparison walks through that path. Real-Time Chunking under --inference.type=rtc has its own execution_horizon setting, default 10, also separate from n_action_steps.
| Situation | Starting n_action_steps | Basis |
|---|---|---|
| SmolVLA on LIBERO | 10 | SmolVLA Table 13; issue 4614 paired runs |
| pi0.5 on LIBERO | 10 | LeRobot LIBERO docs |
| pi0 on a 20 Hz arm | 16 | pi0 paper's 20 Hz setting |
| pi0 on a 50 Hz arm | 25 | pi0 paper's 50 Hz setting |
| ACT, inference fits in one tick | 1 with temporal_ensemble_coeff=0.01 | ACT paper; LeRobot ensembler docstring |
| ACT, inference does not fit | sweep 10 to 50 | ACT docstring example executes 50 of 100 |
| Diffusion, trained at horizon 16 | 8 | Diffusion Policy T_a; diffusion_pusht config |
One call exceeds n_action_steps ticks | raise it, or move to async or RTC | sync path stalls at chunk boundaries |
| Async inference | not used | tune chunk_size_threshold instead |
Write the execution horizon into the checkpoint you ship
The execution horizon changes robot behaviour with identical weights, and it lives in a mutable config.json. Two things can change it after you signed off: an upstream edit to the Hub config you load with --policy.path, and a switch to --policy.pretrained_path that resets it to 50 or 100.
Nor does the evaluation artifact preserve it. In v0.6.1, lerobot-eval writes eval_info.json with per-episode reward, success and seed, and aggregated pc_success, rewards and timings. It does not record n_action_steps, chunk_size or fps. The full resolved config is printed to the log at startup and not written to a file, so a success rate with no log beside it cannot tell you which horizon produced it.
For a regulated cell this is an audit gap, not a convenience issue. Pull the checkpoint into your own registry, set n_action_steps (and temporal_ensemble_coeff for ACT) in the config.json you ship, and record that value in the acceptance evidence next to the success rate and the step id; pinning the checkpoint you deploy covers the step id and digest. If anyone changes the horizon later, re-run the acceptance block, because it is a behaviour change even though no weight moved. Shorter horizons also mean more forward passes per second, which is a sizing question for on-prem inference hardware: at 20 Hz, n_action_steps=10 means two policy calls per second where the default 50 meant 0.4. The rest of the physical AI pillar covers the evaluation and deployment pieces around this one.
This week, take the checkpoint you evaluate most often, open its config.json, and compare n_action_steps with chunk_size. If they are equal, run the four-value sweep above with --output_dir per value and keep the logs with the results.
FAQ
Quick answers to the questions this post tends to raise.



