Contact-rich manipulation depends on force sensing that vision alone cannot provide, both for collecting demonstrations and for training policies. Force-annotated data, however, remains hard to obtain at scale: real-robot collection ties every demonstration to physical hardware, physics simulators report contact forces that deviate systematically from real measurements, and learned world simulators, though scalable and realistic, are vision-only, so operators feel nothing during data collection and the data carries no force/torque (F/T) labels.
We present HapticWorld, an interactive world simulator that predicts joint torque together with observations and renders it back to the operator in real time, closing the haptic loop between a human and a learned world model. Across three contact-rich tasks, torque feedback raises collection throughput by 1.6× on average; policies trained purely on HapticWorld-generated demonstrations succeed in 54/60 real-world trials, approaching the 56/60 of upper-bound real-world data and far exceeding the 19/60 of the vision-only pipeline; and the success rates measured inside HapticWorld closely match real-world evaluation, demonstrating that HapticWorld can serve as a stand-alone F/T-conditioned policy evaluation platform.
Trained on play data, no physics engine. We teleoperate the real robot with no task objective and record camera images, commanded joint positions, and the joint torques the arm already reports. HapticWorld needs no object models, contact parameters, or per-scene system identification.
Torque as a readout of the dynamics model. HapticWorld builds on the visual backbone of the Interactive World Simulator, an action-conditioned consistency model. Instead of a separate torque predictor, a lightweight head reads the widest intermediate features of the dynamics network and maps them to per-joint torque, adding only 181K parameters (about 0.5% of the model). The predicted torque is therefore consistent with the visual future by construction, and the torque loss trains the trunk to carry contact information.
Real-time haptic rendering. Torque comes from the same forward pass that generates the next frame, so the haptic channel adds no inference cost and the loop runs at 10 Hz. Gravity is subtracted, the signal is filtered and clamped, and the interaction torque is applied to all eight axes of the leader arm, including the gripper.
Three visually distinguishable kettlebells of 5, 10, and 15 lb sit on the table. Pushing each produces a markedly different resistance, a direct check that the rendered torque reflects the object being pushed.
Pushing the 10 / 5 / 15 lb kettlebells: real-world operator, HapticWorld view, and rendered force
The operator draws a slingshot and releases a projectile into scoring zones. The rendered draw force and the resulting flight distance grow together with how far the slingshot is pulled, landing the projectile in different zones.
Shot 1: yellow zone (shortest draw)
Shot 2: red zone
Shot 3: blue zone
Shot 4: green zone (longest draw)
We evaluate HapticWorld on three contact-rich tabletop tasks, from data collection to real-robot deployment and in-simulator evaluation.
HapticWorld is built and used in two phases, involving two kinds of data.
Play data trains HapticWorld itself. In the first phase we teleoperate the real robot with no task objective: the operator simply explores the scene while the head-camera images, the commanded joint positions from the leader arm, and the joint torques measured on the follower arm are recorded. For box flipping, for example, play data contains pushing the box in random directions, pivoting it and releasing halfway, or moving the arm without touching the box at all. For each task we collect 500 training and 50 validation play episodes of roughly 20 s each; training HapticWorld takes about two days on eight NVIDIA A100 GPUs.
Demonstration data trains the policies. In the second phase the operator teleoperates inside the trained model with real-time torque feedback to collect goal-oriented task demonstrations, whose observations and torque labels are both generated by HapticWorld. For box flipping, a demonstration is a complete, successful flip. We collect 50 demonstration episodes per configuration and train a torque-conditioned ACT policy on each; these policies are then evaluated either inside HapticWorld or in the real world.
Play data collection on the real robot
The operator drives the follower arm through a synchronized leader arm while the head-camera view, commanded joint positions and measured joint torques are recorded. This is the only phase that touches physical hardware.
We evaluate on three contact-rich tabletop tasks. For each task, play data is task-free exploration of the scene, and demonstration data consists of complete, successful executions collected inside HapticWorld.
Our experiments are designed to answer four questions, which correspond to the main claims of this work:
Evaluation protocol. Each policy is trained once and evaluated over 20 trials with randomized initial configurations, both in the real world and inside HapticWorld. Each rollout is limited to 20 s.
On 50 held-out validation episodes per task, the predicted joint torques closely track the measured ground truth. Overall torque RMSE is 0.38 N·m on microwave opening, 0.26 N·m on whiteboard wiping, and 0.36 N·m on box flipping, i.e. 7–10% of the average contact torque on the most-loaded joint. Use the menu to view the prediction on each task.
Predicted (dashed orange) and measured (solid blue) torque for all eight joints on a held-out validation episode; per-joint RMSE in the panel titles.
Without torque feedback the operator cannot perceive contact, and the applied force drifts outside the range covered by the real play data in either direction: under-application leaves the task incomplete, while over-application pushes the commanded pose into regions the play data never contains and the rollout degrades. Feeling the predicted torque restores the physical constraint inside the model, so collection is both faster and yields better demonstrations.
Demonstrations collected in 10 minutes of teleoperation inside HapticWorld.
| Collection condition | Microwave opening | Whiteboard wiping | Box flipping |
|---|---|---|---|
| With haptic feedback | 26 | 19 | 22 |
| Without haptic feedback | 21 | 8 | 14 |
| Relative throughput | 1.2× | 2.4× | 1.6× |
We collect 50 demonstrations per configuration and train an ACT policy on each. Policies trained purely on HapticWorld demonstrations reach 19/20, 19/20 and 16/20 in the real world, within one trial of policies trained on real-world demonstrations on every task, and far above the complete vision-only pipeline (collecting inside the vision-only simulator and training an image-only policy).
Real-world success rate (out of 20 trials) of ACT policies trained on 50 demonstrations per configuration. Real-world data serves as a reference upper bound.
| Data source | IWS (vision-only) | HapticWorld w/o feedback | Real-world data | HapticWorld (Ours) |
|---|---|---|---|---|
| Task 1 – Microwave opening | 10/20 | 14/20 | 19/20 | 19/20 |
| Task 2 – Whiteboard wiping | 1/20 | 5/20 | 20/20 | 19/20 |
| Task 3 – Box flipping | 8/20 | 6/20 | 17/20 | 16/20 |
Autonomous rollouts side by side
Autonomous rollouts side by side
Autonomous rollouts side by side
Top row: policy without the torque modality, trained on IWS data. Bottom row: torque-conditioned policy trained on HapticWorld data. Left: real-world rollout (inset: head-camera view). Right: the same policy rolled out inside the simulator, as evaluated in Q4. Force traces are shown below the HapticWorld rollouts.
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Vision-only pipeline (IWS)
Trained on real-world data
Trained on HapticWorld data (Ours)
Real-world rollouts of ACT policies trained on 50 demonstrations from each source, shown in real time. Torque traces below the frame show the gripper force observed by the torque-conditioned policies.
Because HapticWorld predicts torque, force-conditioned policies can be rolled out inside it, which a vision-only simulator cannot host. Across all tasks and configurations, in-simulator evaluation preserves the ordering of policies and deviates from real-world success by at most three trials out of twenty.
Policy success rate (out of 20 trials). Each policy is evaluated both in the real world and inside HapticWorld.
| Task | Trained on real-world data | Trained on HapticWorld data | ||
|---|---|---|---|---|
| Eval: Real | Eval: HapticWorld | Eval: Real | Eval: HapticWorld | |
| Task 1 – Microwave opening | 19/20 | 19/20 | 19/20 | 19/20 |
| Task 2 – Whiteboard wiping | 20/20 | 20/20 | 19/20 | 20/20 |
| Task 3 – Box flipping | 17/20 | 14/20 | 16/20 | 14/20 |
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Trained on real-world data, rolled out inside HapticWorld
Trained on HapticWorld data, rolled out inside HapticWorld
Vision-only (IWS) policy, rolled out inside IWS
Policy rollouts inside the simulator: every frame and every torque value is generated by the world model. Compare with the real-world rollouts of the same policies above.
A single, uncut teleoperation episode of more than two minutes inside HapticWorld. Both the predicted video and the predicted torque remain stable over the whole interaction: the scene does not drift or collapse as the autoregressive rollout grows, and the rendered torque keeps tracking contact throughout.
Frames are generated entirely by HapticWorld in response to the operator's commands; no real camera is involved.
We presented HapticWorld, an interactive world simulator that predicts interaction torque jointly with visual observations and renders it back to the operator in real time, closing the haptic loop between a human operator and a learned world model. A lightweight readout of the dynamics network's intermediate features recovers joint torque accurately enough to be felt during teleoperation and consumed by downstream policies.
Across three contact-rich tasks, haptic feedback raises collection throughput by 1.6× on average, and the torque-annotated demonstrations collected inside HapticWorld are sufficient to train torque-conditioned policies that succeed in the real world, approaching policies trained on real-world data and far exceeding the vision-only pipeline. In-simulator evaluation further preserves real-world policy results for this policy class, which vision-only simulators cannot achieve.