Model family and training data
Sparsh is a family of self-supervised learning models for vision-based tactile sensors. The paper reports pretraining on more than 460,000 tactile images, using masking and self-distillation objectives in pixel and latent spaces.
Supported sensors and preprocessing
The reported sensor set covers DIGIT, the marker-based GelSight 2017 and the markerless GelSight Mini. For DIGIT and GelSight Mini images, the researchers subtract a no-contact background; they report that this helps models generalize across sensors of the same type.
Input construction and inference
For image-based self-supervised learning, Sparsh combines two tactile frames taken five samples apart into a six-channel input. At a 60 FPS sensor rate, the paper says that span is about 80 milliseconds. The authors measured inference at up to 112 FPS on an Nvidia RTX 3080 GPU.
Model variants
In the reported TacBench evaluation, Sparsh (DINO) and Sparsh (IJEPA) outperformed Sparsh (MAE). The authors interpret that result as evidence for latent-space learning over pixel-space learning for tactile images.
TacBench results
TacBench contains six touch-centered tasks spanning tactile properties, physical perception and manipulation planning. Under the paper's limited-label setup—33% to 50% of the collected labels depending on the task—Sparsh pretraining improved average performance by 95.1% over task- and sensor-specific end-to-end models.
Sources
This article was researched and fact-checked against the following sources: