AGI Is Not Multimodal: Why Embodiment Matters

Benjamin Spiegel argues that multimodal approaches to AGI will fail, advocating for embodiment and interaction as primary pathways to general intelligence.

MiHiR SEN
MiHiR SEN
·3 min read
Benjamin Spiegel argues that multimodal approaches to AGI will fail because they treat modalities as separate streams rather than emergent phenomena. He advocates for embodiment and interaction as primary pathways to general intelligence, contending that the real challenge is conceptual, not mathematical.

The recent successes of generative AI models have convinced some that AGI is imminent. While these models appear to capture the essence of human intelligence, they defy even our most basic intuitions about it. They have emerged not because they are thoughtful solutions to the problem of intelligence, but because they effectively exploit scale. Seduced by the fruits of scale, some have come to believe that it provides a clear pathway to AGI.

The Multimodal Fallacy

The most emblematic case of this is the multimodal approach, in which massive modular networks are optimized for an array of modalities that, taken together, are supposed to constitute general intelligence. However, this strategy is sure to fail in the near term; it will not lead to human-level AGI that can perform sensorimotor reasoning, motion planning, and social coordination. Instead of trying to glue modalities together into a patchwork AGI, we should pursue approaches to intelligence that treat embodiment and interaction with the environment as primary, and see modality-centered processing as emergent phenomena.

LLMs and World Models

It has been suggested by some that LLMs are learning a model of the world through next token prediction, but it is more likely that LLMs are learning bags of heuristics to predict tokens. This leaves them with a superficial understanding of reality and contributes to false impressions of their intelligence.

One source of evidence in favor of the LLM world modeling hypothesis is the OthelloGPT experiment, wherein researchers were able to predict the board of an Othello game from the hidden states of a transformer model trained on sequences of moves. However, there are issues with generalizing these results to models of natural language. Whereas Othello moves can be used to deduce the full state of an Othello board, we have no reason to believe that a complete picture of the physical world can be inferred by a linguistic description. Othello fundamentally resides in the land of symbols, and is merely implemented using physical tokens to make it easier for humans to play.

The Bitter Lesson and Structure

Sutton's Bitter Lesson has sometimes been interpreted as meaning that making assumptions about the structure of AI is a mistake. This is both unproductive and a misinterpretation; it is precisely when humans think deeply about the structure of intelligence that major advancements occur. Despite this, scale maximalists have implicitly suggested that multimodal models can be a structure-agnostic framework for AGI.

Instead of pre-supposing structure in individual modalities, we should design a setting in which modality-specific processing emerges naturally. What we will lose in efficiency we will gain in flexible cognitive ability. The most challenging mathematical piece of the AGI puzzle has already been solved: the discovery of universal function approximators. What's left is to inventory the functions we need and determine how they ought to be arranged into a coherent whole. This is a conceptual problem, not a mathematical one.