dynamic-mani POLICY DEPLOYMENT

Guides

On this page

Deploy a policy

Updated: 2026-10-01

A completed checkpoint determines the model architecture and the robot, joint, camera, and observation contracts. Use this sequence for the saved UR5e + Sharpa runs:

text
inspect checkpoint → serve model → preview live predictions → deploy control

The preview stage reads live sensors and sends no robot motion commands. The deploy stage sends policy targets through the shared state manager and actuator safety gates. The two stages use the same checkpoint and inference server, but should not run at the same time because the hand feedback port and cameras need a single owner.

#UR5e + Sharpa rollout demos

ACT and tactile ACT support policy rollouts on the UR5e arm with a left Sharpa hand. ACT uses RGB images and measured joint positions. Tactile ACT adds the five Sharpa fingertips' DEFORM maps. Both use temporal ensembling by default and send joint targets through the shared state manager and safety gates. See the policy architecture for their inputs and action-chunk prediction.

These September 30, 2026 clips retain the first half of each supplied export at its existing playback speed. Both are downsampled to 640 × 360 at 15 FPS, with audio retained.

#ACT rollout demo

Download the 77.6-second ACT demo.

#Tactile ACT rollout demo

Download the 27.7-second tactile ACT demo.

#Choose a checkpoint

The six saved runs are under out/best_checkpoints/9_25_tossing_putaozhi_49/:

Use the same policy name with serve, preview, and deploy. --policy checks the architecture in the checkpoint; it cannot change that architecture. For example, --policy dp requires the dp checkpoint, while the tactile ACT checkpoint requires --policy tactile_act.

The commands below use tactile ACT on Sharpa. To run visual ACT, use --policy act with the checkpoint ending in /act in each command. ACT supports --mode sim, --mode dummy, and --mode hw; tactile ACT requires --mode hw and five fresh physical Sharpa DEFORM streams.

bash
CHECKPOINT=out/best_checkpoints/9_25_tossing_putaozhi_49/tactile_act
uv run dmani-policy inspect "$CHECKPOINT"

inspect prints the saved robot identity, camera names and sizes, action dimension, observation rate, and tactile capability. Select a checkpoint that matches the configured rig before opening devices.

#Start the inference server

In the first terminal, from the Dynamic-mani repository:

bash
uv run dmani-policy serve --policy tactile_act --checkpoint "$CHECKPOINT"

The server loads saved normalization and model weights, warms up inference, then listens on tcp://127.0.0.1:5555. It uses CUDA when available; pass --device cpu to select CPU. --host and --port change the server address; pass the matching --endpoint to preview or deploy when using another address. The default server bind is local loopback.

#Preview a live action chunk

In the second terminal, check the sensor-only graph offline, then launch it:

bash
CHECKPOINT=out/best_checkpoints/9_25_tossing_putaozhi_49/tactile_act
uv run dmani-policy preview --policy tactile_act --robot ur5e-sharpa \
  --checkpoint "$CHECKPOINT" --print-dataflow

uv run dmani-policy preview --policy tactile_act --robot ur5e-sharpa

Live preview gets the model contract from the running server. Its optional --checkpoint checks that the server loaded one exact local checkpoint; it is also useful for --print-dataflow when the server is offline. Only serve loads model weights from a checkpoint.

Open the Viser URL printed by the viewer. Click Begin preview to collect a live observation and predict one chunk. The solid mesh shows measured joints; the cyan mesh and amber tool path show predicted motion. Scrub the chunk step, uncheck Play preview to pause the automatic loop, or click Predict a new chunk. Frozen camera images and, for tactile policies, the five DEFORM maps show the input to that prediction. End preview clears the prediction.

Preview opens passive UR5e receive telemetry, Sharpa state/tactile streams, and the rig's configured ZED and wrist cameras. It creates no RTDE control connection or actuator target topic. The Sharpa tactile destination must already point to the configured host IP; preview will not reconfigure it. --camera NAME=SOURCE can replace a configured camera source. RDP preview shows its open-loop proposal; live reactive rollout can select different actions as new tactile frames arrive.

#Deploy through the safety pipeline

Stop preview before starting deployment. Unlike preview, deployment executes the predicted actions. It uses the camera sources in the selected rig configuration for the views required by the checkpoint:

bash
CHECKPOINT=out/best_checkpoints/9_25_tossing_putaozhi_49/tactile_act
uv run dmani-policy deploy --policy tactile_act --robot ur5e-sharpa \
  --checkpoint "$CHECKPOINT" --mode hw

ACT and tactile ACT now use temporal ensembling by default. The deployment controller continually requests action chunks and averages predictions that cover the same 30 Hz action step. The default weight decay is 0.01; use --ensemble-coeff VALUE to change it. --replan-steps N sets the minimum request interval (one step by default). Use --controller async to select the earlier whole-chunk replacement behavior. The saved checkpoint needs no change.

When the camera setup changes, update the rig configuration. --camera NAME=SOURCE remains available for a one-run source override.

Hardware startup checks the cameras, Sharpa SDK and network, and measured feedback before control begins. The operator must confirm startup with y, wait for IDLE, then press a to engage the policy. a pauses or resumes; b returns to idle; c shuts down. The Viser page shows measured joints as solid meshes and current policy targets as translucent cyan meshes. Hardware rollouts are recorded automatically.

For a non-tactile checkpoint, simulation is the default. Replace the policy and checkpoint together; deployment uses the same configured camera sources:

bash
uv run dmani-policy serve --policy dp \
  --checkpoint out/best_checkpoints/9_25_tossing_putaozhi_49/dp

In a separate terminal, after the DP server is ready:

bash
CHECKPOINT=out/best_checkpoints/9_25_tossing_putaozhi_49/dp
uv run dmani-policy deploy --policy dp --robot ur5e-sharpa \
  --checkpoint "$CHECKPOINT" --mode sim

--mode dummy uses a non-connecting sink for non-tactile policies. Tactile DP, tactile ACT, and RDP require --mode hw with five fresh physical Sharpa DEFORM maps; simulated contact maps are not valid replacements. RDP selects its reactive controller by default. --controller rtc is available for compatible epsilon-prediction diffusion checkpoints; see the Policy Learning guide for timing and guidance settings.

#Check a rollout before launch

With --checkpoint and --print-dataflow, deployment prints the graph without opening the cameras, hand, or arm:

bash
CHECKPOINT=out/best_checkpoints/9_25_tossing_putaozhi_49/dp
uv run dmani-policy deploy --policy dp --robot ur5e-sharpa \
  --checkpoint "$CHECKPOINT" --mode sim --print-dataflow

The local checkpoint and server must match exactly when deployment starts. The server and robot processes check the joint order, camera contract, action shape, and model identity. Missing or stale observations and invalid results stop the policy through the existing safety path. See Policy Learning for controller behavior and RDP's reactive action path.