LeRobot's RTC doc says to use real time chunking and async inference together, but in v0.6.1 the async gRPC policy server calls predict_action_chunk with the observation only, so RTC guidance never runs there, and the pull request that would have combined them (PR 3112) closed unmerged on 2026-09-26. You choose one path. With the GPU on or tethered to the robot host, run lerobot-rollout --inference.type=rtc, which already overlaps inference with execution in a background thread and replans when the queue drains to 30 actions. LeRobot sets inference_delay itself as ceil(worst latency x fps): a 120 ms worst call at 30 fps gives 4, and because it is a running max, one slow call raises it for the rest of the process. The code default schedule is LINEAR while the doc recommends EXP, so set EXP explicitly, keep execution_horizon at 10, and keep max_guidance_weight at 10.0 for the 10-step pi0, pi0.5 and SmolVLA. A GPU on a separate server forces the async path: element-wise chunk blending instead of guidance, on an unauthenticated gRPC port. Start by measuring your worst-case policy call and dividing it by your control period.
Real time chunking in LeRobot has a documentation problem worth knowing before you wire up a robot cell. The RTC page in LeRobot v0.6.1 explains that async inference and RTC fix different things (async removes the idle frames while the robot waits for the policy; RTC removes the jerk where one action chunk hands over to the next) and closes its comparison table by telling you to use both together. The shipped async policy server cannot do that. It calls predict_action_chunk(observation) with the observation and nothing else, so the two inputs RTC guidance needs, the inference delay and the leftover of the previous chunk, never reach the policy.
The pull request that would have joined them, PR 3112 ("Distributed Real-Time Chunking"), opened on 2026-03-09 and was closed unmerged on 2026-09-26. Its description states the gap plainly: the RTC implementation does not support a client-server setup, and async inference does not support RTC inpainting. The last change to the policy server file on main, on 2026-09-24, was a type-checking chore. On 2026-10-07 maintainers closed the matching feature request (issue 2562), saying the async_inference module is being removed in favour of a remote inference engine on the rollout contract; that pull request (4836) is still open, so nothing in a release changes yet. So you are not choosing how to combine the two. You are choosing one.
The good news is that the choice is mostly made for you by a physical fact: where the GPU sits. With the GPU at the cell, the rtc backend of lerobot-rollout is the answer, tuned from your worst-case latency; with the GPU on a separate server, the async gRPC path is. Everything below is read from LeRobot v0.6.1, released 2026-08-03 and still the current release, and from the RTC paper, arXiv 2506.07339.
Can you use real time chunking with LeRobot async inference?
Two serving paths exist in v0.6.1, and they differ in more than where the model runs.
The left column is the one the RTC docs actually support, and it already overlaps inference with execution, which is the job the async server exists to do. The right column is what you fall back to when the GPU cannot be on the same machine as the robot driver.
lerobot-rollout --inference.type=rtc | Async gRPC policy server | |
|---|---|---|
| Where inference runs | Same process as the control loop, in a background thread | A separate server process, reached over gRPC |
| Chunk boundary handling | RTC guidance: freeze the actions that will execute, inpaint the rest | Element-wise blend of overlapping timesteps (aggregate_fn_name, default weighted_average) |
| Inference delay | Computed automatically from measured latency | Not passed to the policy |
| Network exposure | None | add_insecure_port, no credentials, host localhost and port 8080 by default |
| Policies | pi0, pi0.5, SmolVLA (plus GR00T, Evo1 and continuous-mode MolmoAct2 in code) | act, smolvla, diffusion, tdmpc, vqbet, pi0, pi05, groot |
| Fits | GPU on the robot host or a tethered workstation | GPU on a separate box |
What the LeRobot rtc inference backend actually does
The rtc backend is asynchronous already, just not across a network. RTCInferenceEngine starts a daemon thread named RTCInference that calls predict_action_chunk and fills a shared ActionQueue. The control loop pops one action per tick at --fps (default 30). The inference thread checks the queue size, and when it is at or below queue_threshold (default 30) it runs the policy; otherwise it sleeps 10 ms and checks again.
This replaces the select_action queue that the sync backend uses. If you tuned n_action_steps on the sync path, that value stops governing replanning here; the mechanics of that queue are covered in LeRobot n_action_steps vs chunk_size.
When a new chunk arrives, ActionQueue.merge in RTC mode replaces the whole queue with the new chunk minus its first d actions, where d is the delay measured for that call. Those first d actions correspond to ticks that already passed while the model was thinking, so they are dropped rather than replayed.
The arithmetic from the source is worth doing once, because it tells you how often the policy actually sees the world. Pi0, pi0.5 and SmolVLA return a 50-action chunk by default. After a merge the queue holds 50 - d. The thread fires again when the queue drains to 30, and during the next call d more actions are consumed. So about 20 actions from each chunk execute, whatever d is, as long as d stays at or below 20. Past that, the merged queue already sits at or under 30, so the next call fires immediately and the cadence becomes one call per d ticks. At 30 fps that is a fresh observation roughly every two thirds of a second.
The failure edge is also in that arithmetic. If a single call takes longer than queue_threshold ticks (30 ticks, or 1 second at 30 fps), the queue empties, send_next_action returns nothing for those ticks, and no new command goes to the robot until the chunk lands.
How LeRobot sets inference_delay from your latency
You do not set inference_delay in lerobot-rollout. RTCConfig has no such field, and there is no flag for it. The backend computes it before each call in rollout/inference/rtc.py:
ceil(latency_tracker.max() / (1 / fps)), which equals ceil(max latency x fps).perf_counter, adds that to the tracker, and passes ceil(new_latency x fps) to merge() as the real delay.Worked through: at --fps=30, a worst observed call of 120 ms gives ceil(0.12 x 30) = ceil(3.6) = 4. The next chunk is generated knowing that 4 actions will execute before it lands, and those 4 are the frozen prefix.
The maximum never forgets
LatencyTracker.max() returns a running maximum since the last reset(). The tracker keeps a deque(maxlen=100) of samples, but that bound only applies to percentile(), which the rtc backend does not call. The tracker lives in the inference thread and is reset only during torch.compile warmup: with --use_torch_compile on, the first max(1, compile_warmup_inferences) calls are discarded (default 2). engine.reset() between episodes clears the queue and the policy, not the tracker. The consequence: one slow call, from a garbage collection pause, a thermally throttled GPU or a cold cache, raises inference_delay for every later chunk until you restart the process. The RTC paper handles this differently. Its Algorithm 1 keeps a bounded buffer of the last 10 delays in its real-world runs and defines delay with a floor, not a ceiling. LeRobot is more conservative on both counts. That is safe in one direction (a delay that is too high only freezes more actions) and costly in the other: a delay that creeps toward execution_horizon shrinks the guided region to nothing, which the next section explains. Three habits follow from it: To get the worst-case number in the first place, measure the control loop the way the sim-to-real latency walkthrough describes, then divide by your control period.
- 1. Warm up before the episode that matters. If you use
--use_torch_compile, the warmup calls are already excluded; without it, the very first calls count toward the max. - 2. Restart the rollout after any known stall, such as a dataset save or a thermal event, rather than carrying the inflated max forward.
- 3. Watch the log for
Indexes diff is not equal to real delay.merge()logs it when the number of actions actually consumed during inference differs from the computed delay. It is the field signal that your latency is jittering against the control period.
LeRobot RTCConfig settings: what to start from
Three surfaces in v0.6.1 disagree on defaults. The rollout code defaults to LINEAR and a horizon of 10. The doc snippet uses EXP and 10. The offline eval_dataset.py script defaults to EXP and 20. Know which one you are reading.
Validation in __post_init__ only rejects max_guidance_weight or debug_maxlen at or below zero. execution_horizon is not checked against anything.
lerobot-rollout \
--strategy.type=base \
--policy.path=<your_org>/<your_pi05_or_smolvla_checkpoint> \
--inference.type=rtc \
--inference.rtc.execution_horizon=10 \
--inference.rtc.max_guidance_weight=10.0 \
--inference.rtc.prefix_attention_schedule=EXP \
--inference.queue_threshold=30 \
--fps=30 \
--robot.type=so101_follower \
--robot.port=/dev/ttyACM0 \
--task="Place the part in the fixture" \
--duration=120 \
--device=cudaexecution_horizon is not the paper's execution horizon
In the RTC paper, the execution horizon s is how many actions of a chunk execute before switching to the next one, and the overlapping remainder is soft-masked. In LeRobot, execution_horizon is the length of the previous-chunk prefix used for guidance. The quantity that plays the paper's s in a LeRobot rollout is set indirectly, by chunk_size - queue_threshold. The paper's finding that performance rises as its execution horizon falls is about s, measured in simulation with the delay fixed at 1. It says nothing about lowering LeRobot's flag. The guidance weights come from get_prefix_weights(start=inference_delay, end=execution_horizon), which first sets start = min(start, end). The first d steps get weight 1, the steps up to execution_horizon decay, and everything after gets 0. The previous leftover is truncated or zero-padded to exactly execution_horizon steps. Put that next to merge(), which discards the first d actions of every new chunk, and the practical rule falls out: execution_horizon has to be larger than your delay. At a delay of 10 or more with the default horizon, the decaying region vanishes and every guided step lands inside the actions that get thrown away. At 30 fps that happens as soon as the running max passes 300 ms. Keep queue_threshold well above execution_horizon plus your expected delay too; the defaults, 30 against 10, already do.
EXP versus LINEAR
Both schedules give weight 1 to the frozen prefix and 0 past the window. LINEAR decays in a straight line between them. EXP multiplies the linear ramp by expm1(lin) / (e - 1), which is below 1 everywhere inside the window, so it sits under LINEAR and lets go of the old chunk faster. The paper's soft mask decays exponentially as well. The code carries a comment saying the default should change to EXP; until it does, pass it explicitly.
max_guidance_weight belongs to a step count
The guidance weight per denoising step is min(c x inv_r2, max_guidance_weight) with c = (1 - tau) / tau, so max_guidance_weight is a clip, the paper's beta. The docs call 10.0 optimal for 10-step flow matching, and pi0, pi0.5 (num_inference_steps) and SmolVLA (num_steps) all default to 10. The paper used a clip of 5 with 5 denoising steps, chose it as a conservative value from a simulated ablation, and reported that unclipped weights made chunks diverge at low step counts. If you cut denoising steps to buy latency, the 10.0 advice no longer has the docs behind it.
The deployment command
With the GPU on the robot host, this is the starting configuration. EXP is set explicitly because the code default is LINEAR, and --inference.queue_threshold is shown at its default so the knob is visible. The robot type, port and task are placeholders; add your --robot.cameras dictionary as in the LeRobot RTC example. If the policy type does not report RTC support, rollout stops at startup with a ValueError that tells you to use --inference.type=sync instead. The docs list Pi0, Pi0.5 and SmolVLA; the code also accepts GR00T, which maps the leftover into its own native action-overlap options, Evo1, and MolmoAct2 when its inference_action_mode is continuous. If you are still choosing the model, our VLA comparison of OpenVLA, pi0, SmolVLA and GR00T N1 covers memory and fit.
Field (--inference.rtc.<field>) | v0.6.1 code default | Doc advice | What it trades |
|---|---|---|---|
enabled | True | Off means merge appends instead of replacing | |
execution_horizon | 10 | Typical 8 to 12 | Higher is smoother, lower more reactive. Must exceed your delay |
max_guidance_weight | 10.0 | 10.0 for 10-step flow matching | Clip on the per-step guidance weight |
prefix_attention_schedule | LINEAR | EXP to start | How fast guidance decays across the window |
queue_threshold (--inference.queue_threshold) | 30 | Replan trigger; sets how many actions per chunk execute | |
debug, debug_maxlen | False, 100 | Tracing |
LeRobot async inference when the GPU is on another machine
The async server exists for the case where the robot host has no usable GPU. The robot client streams observations, the server returns chunks, and the client decides when to ask for more. What you give up is RTC guidance. Chunk boundaries are smoothed by the client's aggregate_fn_name, which blends overlapping timesteps element by element:
That is averaging two independently sampled chunks, not generating the new chunk to agree with the old one. For a multimodal policy, two chunks can choose different ways around an obstacle, and their average may be neither.
Memory on the server matters if you share one GPU across cells: the async docs put PI0 at about 14 GB at inference time and SmolVLA at about 2 GB. The server's thresholds, client commands and their documentation quirks are covered in the sim-to-real post; they do not change here. The async constants also define a SUPPORTED_ROBOTS list, but the client's check against it is commented out in v0.6.1, so it is not enforced.
aggregate_fn_name | Blend (old, new) |
|---|---|
weighted_average (default) | 0.3, 0.7 |
average | 0.5, 0.5 |
conservative | 0.7, 0.3 |
latest_only | new chunk only |
RTC vs async inference: choosing by GPU location and latency
Compare your worst-case call against the control period (33 ms at 30 fps). No row below carries a measured rate or success figure; the starting settings come from the v0.6.1 source and docs.
The rows that matter are the middle two. A workstation next to the cell, with the robot's serial cable and cameras plugged straight into it, counts as "same host" for this purpose. It runs the whole rollout process, so it gets RTC with no network hop.
| GPU location | Worst-case call vs control period | Path | Start from |
|---|---|---|---|
| Onboard or same host as the robot | Fits inside one tick (rare for 10-step VLAs at 30 fps) | sync (default) | Defaults; tune n_action_steps |
| Onboard or same host | Spans several ticks | lerobot-rollout --inference.type=rtc | execution_horizon 10, EXP, max_guidance_weight 10.0, queue_threshold 30; keep the computed delay well under 10 |
| Tethered workstation running the rollout process, robot on USB or serial | Spans several ticks | rtc, same process, no network | As above; warm up first if you enable --use_torch_compile |
| Separate GPU server reached over the network | Any | async gRPC, no RTC guidance | weighted_average; thresholds per the sim-to-real post; bind only to a segmented cell interface |
Running RTC on-prem: GPU placement in the robot cell
For a plant or a regulated lab that cannot send camera frames off-site, the RTC paper's own setup is the useful reference. Its real-world inference ran on a single NVIDIA RTX 4090 in a workstation in the same building as the robots. Its end-to-end breakdown for RTC was 108.76 ms on the non-mobile setup, of which the network was 6.89 ms, against 138.98 ms on the mobile setup, where the network was 21.20 ms and image resizing 11.22 ms. In both, the robot computer reached the GPU workstation over wired Ethernet on the same LAN. Even with the model dominating, a network hop is a line item paid on every chunk request, and only the async path pays it.
The security picture separates the paths further. The async server is a grpc.server with four worker threads bound by add_insecure_port, with no credentials and no interceptors, and the client connects over an insecure channel. It defaults to localhost, which is safe. Moving the GPU to a shared server means binding a reachable interface on an unauthenticated port that sends commands to a robot arm. In an environment where a change to that network path goes through a security review, the in-process rtc path is the one with nothing to review.
So the recommendation for a regulated deployment is concrete. Put a GPU at the cell, on the robot host or a tethered workstation, and run lerobot-rollout --inference.type=rtc. If a shared GPU server is unavoidable, accept async without RTC guidance, put the cell on its own segmented network, and bind the policy server only to that interface. For staging weights and datasets inside a network with no Hub access, see running LeRobot offline in an air-gapped plant, and for the wider set of decisions about running physical AI inside your own perimeter, the physical AI guide.
There is a cost on the rtc side, too. Guidance backpropagates through each denoising step: the paper measured 97 ms per RTC call against 76 ms for plain pi0.5 on its 4090 with 5 steps, about 21 ms of overhead. LeRobot's policies default to 10 steps, so treat that as the shape of the cost, not your number. Size the cell GPU for the guided call, not the plain one.
Testing RTC offline with eval_dataset.py, then on the robot
Before touching hardware, run the RTC visualiser on recorded episodes with the delay you computed. examples/rtc/eval_dataset.py plots standard denoising against RTC denoising on dataset samples. Pass --rtc.execution_horizon explicitly, because the script's own default is 20, not the rollout's 10. The --inference_delay here is the one place you set the delay by hand; 4 matches the 30 fps, 120 ms example.
python examples/rtc/eval_dataset.py \
--policy.path=<your_org>/<your_checkpoint> \
--dataset.repo_id=<your_org>/<your_dataset> \
--rtc.execution_horizon=10 \
--rtc.max_guidance_weight=10.0 \
--rtc.prefix_attention_schedule=EXP \
--inference_delay=4 \
--device=cudaThen the week's plan, in order:
ceil(worst x fps). If it is 10 or more at the default horizon, fix latency before tuning anything else.lerobot-rollout with the command above, warmed up, and grep the log for Indexes diff is not equal to real delay.The training-side alternative, arXiv 2512.05964, simulates delay during training and conditions on action prefixes directly, removing the inference-time overhead; on real box-building and espresso tasks it matched inference-time RTC while being cheaper to run. It is not in LeRobot v0.6.1. The issue that requested it was closed in favour of a prototype pull request (2830) that is still open, and a separate pull request adding opt-in training-time RTC to pi0.5 (4056) was merged to main on 2026-08-21, after the v0.6.1 release, so it is not in any release yet. For now, the rtc backend with a GPU at the cell is the path that does what the docs promise.
FAQ
Quick answers to the questions this post tends to raise.



