Knowing a world includes learning what is possible within it.
Experience connects action to consequence. John Dewey located the value of experience in the connection between what we do and what happens as a result. Activity becomes instructive when its consequences inform what follows. For Curio, this inspires a practical question: how can each recorded interaction make the next choice more informed? [5]
A body gives the world possibilities. Embodied approaches to cognition examine how bodies and their relations with the environment shape thought and perception. Our engineering analogy is concrete: what a robot can reach, the contact it can make, and the consequences of that contact all matter to the experience it learns from. [6]
A goal needs a way into action. The philosophical discussion of knowing-how, shaped by Gilbert Ryle, asks how practical ability relates to knowledge of facts. For Soma, the useful question is how a language instruction can call on learned physical capabilities. [7]
The child’s play has interests and purposes of its own. Our inspiration is the richness of those encounters. Whether similarly broad experience transfers to later robot tasks is an experimental question. These philosophical traditions motivate our research; they do not establish its results.
Language models offer a powerful computational precedent. A model can learn structure from a large body of text before being adapted to particular tasks. That shared foundation can support many different uses. [1]
Pretraining can also prepare a model to respond to tasks described in language or illustrated with examples in its context. The final request draws on a much longer history of learning. [2]
That is the bottom-up principle we want to investigate in robotics: let reusable capabilities grow from broad experience, then connect them to specific goals.
The shared idea is reusable preparation. For language, broad text supports a pretrained model that can be prompted or adapted for many tasks. For robotics, actions and their consequences form the experience; language grounding and post-training connect that foundation to useful action.
Fine-tuning refines and adapts a pretrained foundation. The diagram is a conceptual analogy. It does not imply that every capability first appears during fine-tuning, that Soma has collected a complete physical distribution, or that these target skills have been demonstrated. Coverage and transfer remain to be measured.
But a robot’s experience has a different shape. Contact can slip. A drawer can resist. The same movement can produce different outcomes as the scene changes. A foundation for action needs experience of those consequences.
And every physical interaction has a cost. Someone—or something—must decide what happens next.
This turns pretraining into a second problem: how do we acquire the experience worth learning from?