75%
Average success
across five tasks
+43 pts
over no tactile input
32% → 75%
+7 pts
over tactile input alone
68% → 75%
See it in action

Uni-VLaT in the Real World

Four real-robot rollouts: Table Sweeping, Basket Loading, Back-Tap Walking, and Human–Robot Hugging.
Five contact-rich tasks

Real-World Tasks

From contact-triggered walking to sustained support and sequential cleanup. Watch each real-robot rollout below.

Table SweepingContact-guided object sweeping
Basket LoadingSupport through changing load
Back-Tap WalkingMove on contact, stop on release
Human–Robot HuggingResponsive embrace and release
Composed CleanupSequential gather, carry, and deposit
The missing signal

Why Touch?

Vision can miss occluded contact. Proprioception shows how the robot moves, but only indirectly reveals what it touches. Distributed tactile sensing shows where contact happens across the body and how it changes.

Uni-VLaT combines whole-body touch with vision, language, and proprioception for humanoid loco-manipulation.
Whole-body tactile sensing complements visual, language, and body-state inputs across the five tasks.
How it works

Method

Spatial and temporal tactile encoding feeds a pretrained VLA policy. Training-only heads predict future tactile, proprioceptive, and visual representations.
The predictive heads train interaction-aware representations; they are omitted at deployment.
01

Whole-body tactile encoding

Regional sensor readings become spatial tactile tokens; a short causal history captures contact onset, support, and release.

02

Tactile-anchored multimodal prediction

Contextualized tactile features predict future tactile, proprioceptive, and visual representations during training.

03

VLA adaptation

The tactile pathway adapts pretrained VLA policies for whole-body actions while retaining their visual and language priors.

Unitree G1 · Isaac-GR00T

Main Results

Average success rises from 32% without tactile input to 75% with Uni-VLaT. Each task uses 50 demonstrations; the main configurations use 20 real-robot rollouts per task.

Back-Tap Walking is primarily enabled by tactile input: tactile-only prediction reaches 90%, versus 85% for full Uni-VLaT.

More than one policy

Generalizes Across VLA Backbones

Uni-VLaT improves success on Back-Tap Walking, Basket Loading, and Table Sweeping with both Isaac-GR00T and π0.5. It also raises the normalized stage score on Composed Cleanup.

Isaac-GR00TSweeping 45% → 75%Walking 0% → 85%
π0.5Sweeping 30% → 60%Walking 0% → 90%
Cross-backbone results. With Uni-VLaT, Isaac-GR00T reaches 85% on Back-Tap Walking, 80% on Basket Loading, and 75% on Table Sweeping; pi 0.5 reaches 90%, 60%, and 60%. Composed Cleanup normalized stage scores rise from 43.3% to 70% and from 26.7% to 46.7%, respectively.
(a) Real-robot success rates on three tasks; (b) normalized stage score on Composed Cleanup. Isaac-GR00T uses 20 rollouts per configuration and π0.5 uses 10. View the vector PDF.
π0.5 · Table Sweeping
π0.5 · Back-Tap Walking
Contact in motion

Touch Reveals Interaction Dynamics

During Basket Loading, tactile feedback changes as objects are added and the basket is removed. The curve shows the right-arm response around these events.

No Tactile · Basket Loading rollout
Uni-VLaT · Basket Loading rollout
Basket Loading contact-response curves: Uni-VLaT changes with successive object placements while No Tactile remains at a relatively high level.
Right-arm tactile response around loading events

The trace is from a recorded rollout, in arbitrary ADC units; it is not a calibrated force or aggregate safety measure. No Tactile sensor readings were recorded for analysis only and were not policy inputs.

What makes it work?

Ablations

Average success on Table Sweeping and Back-Tap Walking shows the value of post-policy tactile context and absolute future targets.

Context: Post-DiT tactile states outperform earlier or generic multimodal contexts.

Targets: Predicting absolute future latents preserves sustained contact information better than predicting changes alone.

Full Uni-VLaT reuses the 20-rollout main evaluation; other ablations use 10 rollouts per task.

Paper summary

Abstract

Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception provide only indirect evidence of interaction, especially when the contact region is occluded. Distributed tactile sensing preserves spatially resolved contact patterns across the robot body.

Uni-VLaT adapts pretrained vision-language-action policies with a tactile pathway trained for both action generation and prediction of future tactile, proprioceptive, and visual representations. Contextualized tactile features anchor these complementary views of physical interaction. Across five real-robot tasks, Uni-VLaT achieves 75% average success, compared with 32% for No Tactile and 68% for tactile input without prediction. Evaluation with Isaac-GR00T and π0.5, together with controlled ablations, supports the benefits of post-DiT tactile context and absolute future targets.

Cite this work

BibTeX

Preliminary citation; paper and arXiv identifiers will be added when released.

@misc{wang2026univlat,
  title  = {Uni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-Manipulation},
  author = {Wang, Zihao and Liu, Shutong and Zheng, Siqi and Cao, Liu and Chen, Ruoqi and Liu, Rundong and Yang, Yanchao and Xu, Mengdi},
  year   = {2026},
  note   = {Project page: https://ggkiller-air.github.io/Uni-VLaT/}
}