Image: thegradient.pub · rights & removal
AGI Is Not Multimodal
Reporting by The Gradient AI PublicationRead the original at thegradient.pub
Executive Summary
Facts Only
* Terry Winograd stated that projecting language back as the model for thought results in losing sight of tacit embodied understanding.
* Generative AI successes do not indicate imminent AGI because models emerged from scaling on existing hardware rather than thoughtful solutions to intelligence problems.
* Multimodal approaches, involving massive modular networks optimized for various modalities, are argued to fail in the near term for achieving human-level AGI that includes sensorimotor reasoning and motion planning.
* True AGI requires an ability to solve problems originating in physical reality, such as repairing a car or preparing food.
* The predict-next-token objective is suggested to yield bags of heuristics rather than true world models.
* LLMs' performance suggests they induce world models via next-token prediction, supported by observations like the Othello paper.
* A full physical world model is required for solving many problems that cannot be converted into symbol manipulation.
* The reasoning behind LLM behavior might stem from brute force memorization of syntax rather than a learned world model.
* Syntax studies sentence structure; semantics deals with literal meaning (information about the world); pragmatics involves contextual interaction and environment reasoning.
Full Take
From the original · The Gradient AI Publication
"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry Winograd The recent successes of generative AI models have convinced some that AGI is imminent. While these models appear to capture the essence of human intelligence, they defy even our most basic intuitions about it.Read the full story at thegradient.pub
