Visual TL;DR. VLA Scaling Bottleneck leads to Conflated Learning. Conflated Learning reveals Decomposition Hypothesis. Decomposition Hypothesis enables TAP Framework. TAP Framework includes Stage 1: Motor Priors. TAP Framework includes Stage 2: Language Grounding. Stage 1: Motor Priors results in Efficiency Gains. Stage 2: Language Grounding contributes to Efficiency Gains. Stage 1: Motor Priors results in Robustness Gains. Stage 2: Language Grounding contributes to Robustness Gains.
- VLA Scaling Bottleneck: prohibitive cost of collecting expert demonstrations for VLA models
- Conflated Learning: physical competence and semantic alignment learned together
- Decomposition Hypothesis: physical competence needs no language supervision
- TAP Framework: task-agnostic pretraining for embodied AI
- Stage 1: Motor Priors: learns transferable motor priors from unlabeled data
- Stage 2: Language Grounding: grounds physical representations with minimal expert data
- Efficiency Gains: orders of magnitude efficiency with minimal labeled data
- Robustness Gains: demonstrates superior robustness on downstream tasks
Visual TL;DR
