N0-TWAM now has a public code repository and downloadable model bundles. The release includes an inference server and tools for post-training on new demonstrations, but it does not include the large-scale pretraining pipeline.

What the release contains

The repository provides code for loading the pretrained model, adapting it to a new robot, serving actions over a websocket and running closed-loop evaluation in NeoSim. Its example client is designed to connect a robot or simulator to that server.

Five model bundles are linked from the repository: one pretrained base model, two post-trained UniVTAC variants and two post-trained NeoSim variants. The task-specific bundles separate absolute end-effector actions from horizon-delta actions.

The code and model are released under a noncommercial Creative Commons share-alike license.

How N0-TWAM predicts an action

N0-TWAM is a world-action model built around video, tactile and action streams. It first predicts the coming visual scene and contact signal, then generates an action conditioned on that predicted future.

The architecture assigns private transformer capacity to each stream while allowing the three streams to share attention. The project page lists a full-width video expert alongside slimmer tactile and action experts, for a reported total of 7.16 billion trainable parameters.

Touch enters the system in two different forms. One path predicts future tactile video in the same latent space used for scene video. A second path converts current InTac S1 sensor images into three-axis force maps and then into NeoForce tokens that condition the action head.

The training sequence keeps those roles separate. Large-scale pretraining uses predicted touch, while task-specific post-training activates the observed-touch input and fine-tunes NeoForce with the action objective.

Training data and reported tests

The project says NeoData spans six robot embodiments and 450 tasks, producing about 7.5 million aligned training clips. Those clips combine camera views, per-finger tactile video, language and a 20-dimensional bimanual end-effector action.

For evaluation, the authors used eight UniVTAC tasks, twelve NeoSim tasks and eight contact-rich tasks on physical robots. They report 100 randomized trials for each simulation task and 20 trials for each physical-robot task.

The reported average success rates are 84.5% on UniVTAC, 49.4% on NeoSim and 46.3% on the physical-robot tasks. The figures should therefore be read as team-reported benchmark results.

What researchers can reproduce now

The public materials are enough to inspect the model implementation, download the base weights, run the inference server and use the documented post-training path. They also include post-trained bundles for the two simulation suites.

The release does not provide the data or pipeline used for large-scale pretraining. Reproducing the complete training process therefore requires more than the published repository, even though loading, serving and task-specific adaptation are documented.

Sources

This article was researched and fact-checked against the following sources: