TACO's authors report an average task score of 0.825 after two post-training rounds, versus 0.375 for the base policy. The associated repository says its current release omits the IDVM, the full correction pipeline, checkpoints and datasets.
What TACO does
TACO first uses its IDVM to flag states near a failure. A separate model proposes local corrections as paired visual and tactile sequences. The IDVM then supplies actions for correction and estimates of progress.
Before ranking a correction on predicted progress gain, TACO screens whether its motion is feasible and its touch signal is plausible.
What the reported result measures
The repository says the evaluation spans six real-world tasks and uses 40 episodes for each task. Its table moves from a 0.375 base-policy task score to 0.650 after one TACO iteration and 0.825 after the second. The authors define task score to include partial completion; their completion-step metric covers successful episodes only.
What is available in the repository
According to the current README, code is staged for two modules. One generates visuo-tactile sequences; the other supplies the tactile VLA policy.
For the generation side, the repository provides denoising across video and touch, loaders for force sequences, runnable training and caching paths, inference code and sample configurations.
The policy side uses a pi0.5-style flow-matching path with an eight-step force history. Its conditioning uses binary advantage labels, and it also supports classifier-free guidance. Separate instructions cover conversion of data as well as training and inference.
What is still missing
The repository explicitly says the inverse-dynamics and value model is not included. Checkpoints and datasets are also absent, along with code for the complete iterative correction pipeline.
Sources
This article was researched and fact-checked against the following sources: