AI models flub these intelligence tests. Can you fare any better?
AI models still flunk basic visual and logic puzzles where humans routinely succeed, exposing stubborn gaps in machine cognition even after rapid advances last year.
Source: MIT Technology Review · August 26, 2026 at 12:01 PM · AI-assisted report
Single-sourceKUALA LUMPUR, MALAYSIA, NEW YORK TIMES, UNIVERSITY OF ILLINOIS URBANA-CHAMPAIGN, GOOGLE, 26 AUGUST 2026 —
AI models still flunk basic visual and logic puzzles where humans routinely succeed, exposing stubborn gaps in machine cognition even after rapid advances last year.
Market Impact
Columbia University researchers showed in December 2024 that top large language models solved only 18% of New York Times Connections puzzles, but by March 2025 some approached near-perfect accuracy. The same work found models still stumble when wording shifts slightly or visual cues change, despite vast training data.
Spatial reasoning remains a blind spot. A University of Edinburgh study published in 2025 found vision-language models perform near chance on mental rotation tasks used in IQ tests, where humans typically score above 80%. The paper tested whether machines could identify the same 3D object from different angles—an ability architects and engineers rely on daily.
Classic logic puzzles also reveal flaws. In a 2024 experiment by Google and the University of Illinois Urbana-Champaign, models trained on variations of the Knights and Knaves riddle often repeated memorized solutions rather than adapting to new constraints. “Models depend on surface patterns rather than genuine reasoning,” the authors wrote in Nature Machine Intelligence.
Scale magnifies the problem. Apple researchers reported in February 2025 that large language models solve simple Towers of Hanoi and river-crossing puzzles flawlessly with small numbers of disks or travelers. Once complexity rises—six disks or six travelers—accuracy drops below 40%. A separate study by the University of Washington, Stanford University, and the Allen Institute for AI showed similar breakdowns on logic grid puzzles as the number of attributes and people increased.
For Malaysian businesses, the findings explain why AI cannot yet replace human intuition in areas such as design, diagnostics, or strategic planning where spatial manipulation and adaptive reasoning matter. While models may soon master trivia and text-based benchmarks, tasks requiring spatial ability or robust deductive logic remain beyond their reach.
Related: Malaysia Digital Economy Corporation (MDEC)