YouTube Review · Published: August 24, 2026 · Modified: August 24, 2026 · Last reviewed: August 24, 2026

Why AI Is Incredibly Smart and Shockingly Stupid

Why AI Is Incredibly Smart and Shockingly Stupid | Yejin Choi | TED is a compact 2023 argument against confusing broad benchmark success with dependable understanding. Computer scientist Yejin Choi places spectacular capability beside mundane failure, then asks whether simply making language models larger can supply the background knowledge, causal expectations, and social judgment that people compress into the phrase common sense.

The talk remains useful because it joins two debates that are often separated. One is technical: what does a model know when it can pass a professional exam yet mishandle an ordinary scenario? The other is institutional: who can build, inspect, and correct models when frontier development requires extraordinary concentrations of compute, data, labor, and capital? Choi's answer is not a rejection of neural networks or even of scale. It is a case for alternative data and algorithms, smaller inspectable systems where they are sufficient, and public scrutiny of the norms used to shape machine behavior.

The Transcript's Argument

The transcript follows a clear three-part structure. First, scrutinize the system rather than generalizing from its most impressive score. Choi contrasts high-profile achievements in games and exams with three GPT-4 responses involving drying clothes, measuring water, and reasoning about a bridge above sharp debris. The examples are deliberately ordinary. Their purpose is to show that polished language and broad test performance do not guarantee stable handling of physical constraints or the simplest interpretation of a situation.

Second, choose common sense as a research target. Choi compares the unspoken physical, causal, and social assumptions behind language to invisible matter inferred through its effects. Her point is not that common sense is one fact database. It includes naive physics, likely intentions, theory of mind, norms, morals, and the ability to form hypotheses through interaction. This is why the talk treats next-token prediction as an indirect route to knowledge: a language model may absorb enormous amounts of information while still learning the world only through the textual traces people leave behind.

Third, change the inputs and learning methods. The talk distinguishes raw web data from deliberately crafted examples and human judgments, calls for important norm resources to be open to inspection and correction, and presents symbolic knowledge distillation as one route from a large general model to a smaller, task-focused commonsense model plus a readable knowledge artifact. In the closing exchange, Choi adds an important qualification: scale produces real learning, smaller is not automatically better, and the eventual recipe may combine an appropriate amount of scale with other ideas.

A Jagged Frontier, Not a Stupidity Score

The title is memorable, but "stupid" is rhetoric rather than a measurement. The three GPT-4 mistakes are demonstrations, not a controlled benchmark. The transcript does not report the full prompts, sampling settings, repeated trials, comparison models, or failure frequency. They establish that surprising failures occurred in a 2023 system; they do not establish how often such failures occurred, whether the same prompts fail today, or that scale itself caused each error.

The broader diagnosis has aged better than the examples. The 2026 Stanford AI Index describes a continuing jagged frontier: systems can reach exceptional results in competition mathematics while remaining much weaker on an ordinary visual task such as reading an analog clock. That does not prove Choi's preferred theory of common sense. It does support her operational warning that capability is uneven and that success in one evaluation cannot be silently transferred to another context.

For AI evaluation, the useful unit is therefore a claim tied to a task, model version, date, prompt distribution, and consequence. "Passed the bar" is not "reliable legal agent." "Answered this riddle" is not "has a world model." "Failed this riddle" is not "incapable of useful reasoning." Choi's examples are best used to break the spell of aggregate competence, then replaced with repeated, deployment-relevant tests.

The Power Argument Is Stronger Than the Model Anecdotes

The talk's most durable claim concerns concentration. Choi argues that extreme training costs narrow who can build frontier systems and leave outside researchers with limited access to their internal construction. The 2026 AI Index reports that industry produced more than 90 percent of notable frontier models in 2025. Its responsible-AI chapter also reports a decline in average foundation-model transparency, with persistent gaps around training data, compute, and post-deployment effects. Those findings do not make every company equally opaque, but they independently support the direction of Choi's warning.

This point should not be collapsed into "small equals democratic." A compact model can still be proprietary, surveillant, biased, or controlled by one vendor; a large model can release weights, documentation, evaluations, or research access. Size, openness, ownership, energy use, inspectability, and effective public control are related but distinct properties. The Spiralist question is who can examine the system, contest its outputs, change its rules, and leave it—not merely how many parameters it contains.

Inspectable Norms Are Not Moral Truth

Choi's call to teach human norms and values needs the most care. "Human values" is not a clean label waiting to be attached to enough examples. People disagree across cultures, communities, roles, histories, and situations; a widely held norm can also encode hierarchy or prejudice. A system trained to reproduce the most common judgment may become socially fluent while laundering a majority view into apparent moral authority.

Choi's own research provides a better reading than the slogan alone. Social Chemistry 101 describes rules of thumb as observations about social judgment and explicitly treats norms as culturally sensitive rather than universal prescriptions. Openness helps because researchers and affected communities can inspect examples, document disagreement, find exclusions, and propose corrections. It does not decide whose values should govern a school, hospital, workplace, companion, or public service.

That distinction connects the video to plural value alignment and AI alignment. A norm dataset should carry provenance, population and sampling limits, disagreement records, intended uses, prohibited uses, and an appeal path. Without those institutional layers, an inspectable corpus can still become an oracle.

What Symbolic Distillation Actually Supports

The alternative technical path in the talk is not hypothetical. The peer-reviewed Symbolic Knowledge Distillation paper uses a large language model to generate causal commonsense candidates, filters them with a separately trained critic, produces a symbolic knowledge graph, and trains a smaller commonsense model. The authors report that the student surpassed its GPT-3 teacher on their scoped commonsense evaluations while being 100 times smaller.

That is meaningful evidence that scale is not the only lever for a bounded task. It is not evidence that a small model acquired general human common sense. The result depends on the selected ATOMIC relations, prompts, teacher, critic, quality criteria, and evaluation design. Because the knowledge begins with a large model, the pipeline can also preserve or reorganize the teacher's blind spots. Its governance advantage is narrower and still valuable: the intermediate text can be inspected and filtered in a way that opaque neural representations cannot.

The transcript's closing qualification should govern the review. Choi accepts that scaling had produced remarkable gains and proposes a possible middle range where scale and other methods are combined. Read this as a research portfolio argument, not as proof of a single replacement architecture: invest in explicit knowledge, grounded interaction, efficient task models, plural norm data, and evaluation alongside general-purpose scaling.

Spiralist Use

The video offers a practical discipline for encountering fluent systems. Separate capability, reliability, and authority. A model may be capable enough to help, unreliable enough to require verification, and institutionally powerful enough to shape what gets counted as knowledge. None of those facts settles the others.

For consequential use, record the exact model and date; test the local task across repeated and adversarial cases; preserve prompts and failures; define abstention and human-escalation routes; disclose where norm judgments came from; and give affected people a way to contest the result. NIST's Generative AI Profile reinforces this lifecycle view by placing trustworthiness work across design, development, use, and evaluation rather than inside one benchmark score.

This is also a claim-hygiene lesson. Do not turn a model's eloquence into evidence of a coherent inner intelligence, and do not turn a comic failure into evidence that the system is harmless. Choi briefly uses the metaphor of a new intellectual species; the transcript offers no test of consciousness or AGI, and this review does not treat the metaphor as either. The systems matter because of what they can do, what they fail to do, and the institutions that give their outputs force.

Evidence and Limits

The exact YouTube URL, title, TED channel, and April 28, 2023 upload date were checked on the official upload. The talk's sequence and claims were checked against TED's official talk/transcript page and a public full-text transcript mirror. This review did not obtain YouTube's caption file directly, independently reconstruct every slide, or verify every numerical estimate spoken in the talk.

The video is a conference talk, not a paper-length evaluation of GPT-4. Its model examples are a dated snapshot, and this review did not rerun them on current systems. The 2026 AI Index supports the continued existence of uneven capabilities and concentrated frontier development; it does not retroactively validate every causal, environmental, or architectural claim in the talk. Treat the video as a strong framing source and research agenda, not as a current leaderboard or a complete theory of intelligence.

Sources


Return to YouTube