Skip to content
Breaking
XMUM Students Win Bronze at 12th China Undergraduate Medical Innovation TournamentOver 380 missing, including tourists, after flash flood on Nepal-Tibet border kills at least 19IJM 1Q profit jumps 44% on construction boost, declares 10 sen special dividendHearing tech startup Legato emerges from stealth with $12M and a preview of its AI hearing glassesCME Student Among Top Five at BEM Young Engineering Innovation Trophy 2026Man who made more than 1,000 silent calls to police 'for fun and enjoyment' gets jailIJM names Lee Teck Yuen as chairman to succeed Krishnan TanRunable raises $21M to bet AI agents can go from building businesses to growing themAI models flub these intelligence tests. Can you fare any better?Malaysian firms warned against rushing AI model choices for security operationsMudslide from Nepal kills several at Gyirong Port trade hubUS Cold War aid turned Thailand into a regional growth hubApple’s new cleaning cloth now cheaper at RM39Malaysia’s 2026 GDP growth forecast raised to 5.4%Rupiah slips to Rp17,725 as Fed policy clues weighJakarta stocks plunge 1.48% as political transition and risks weighApple Mac Studio 2026 Malaysia: M5 Max and M5 Ultra, priced from RM10,999Sri Aman Highest, Here Are 17 IPU Unhealthy AreasSamsung’s Galaxy Z Series roadshow in Kuala Lumpur draws crowdsIsmail Sabri to face corruption charges tomorrowXMUM Students Win Bronze at 12th China Undergraduate Medical Innovation TournamentOver 380 missing, including tourists, after flash flood on Nepal-Tibet border kills at least 19IJM 1Q profit jumps 44% on construction boost, declares 10 sen special dividendHearing tech startup Legato emerges from stealth with $12M and a preview of its AI hearing glassesCME Student Among Top Five at BEM Young Engineering Innovation Trophy 2026Man who made more than 1,000 silent calls to police 'for fun and enjoyment' gets jailIJM names Lee Teck Yuen as chairman to succeed Krishnan TanRunable raises $21M to bet AI agents can go from building businesses to growing themAI models flub these intelligence tests. Can you fare any better?Malaysian firms warned against rushing AI model choices for security operationsMudslide from Nepal kills several at Gyirong Port trade hubUS Cold War aid turned Thailand into a regional growth hubApple’s new cleaning cloth now cheaper at RM39Malaysia’s 2026 GDP growth forecast raised to 5.4%Rupiah slips to Rp17,725 as Fed policy clues weighJakarta stocks plunge 1.48% as political transition and risks weighApple Mac Studio 2026 Malaysia: M5 Max and M5 Ultra, priced from RM10,999Sri Aman Highest, Here Are 17 IPU Unhealthy AreasSamsung’s Galaxy Z Series roadshow in Kuala Lumpur draws crowdsIsmail Sabri to face corruption charges tomorrow
AI Edge

AI models flub these intelligence tests. Can you fare any better?

AI models still flunk basic visual and logic puzzles where humans routinely succeed, exposing stubborn gaps in machine cognition even after rapid advances last year.

Source: MIT Technology Review · August 26, 2026 at 12:01 PM · AI-assisted report

Single-source
AI models flub these intelligence tests. Can you fare any better?
Photo: Nikita.smagin / CC BY-SA 3.0

KUALA LUMPUR, MALAYSIA, NEW YORK TIMES, UNIVERSITY OF ILLINOIS URBANA-CHAMPAIGN, GOOGLE, 26 AUGUST 2026 —

Listen to this article

DomainFork Audio · read aloud

Share

AI models still flunk basic visual and logic puzzles where humans routinely succeed, exposing stubborn gaps in machine cognition even after rapid advances last year.

Market Impact

Columbia University researchers showed in December 2024 that top large language models solved only 18% of New York Times Connections puzzles, but by March 2025 some approached near-perfect accuracy. The same work found models still stumble when wording shifts slightly or visual cues change, despite vast training data.

Spatial reasoning remains a blind spot. A University of Edinburgh study published in 2025 found vision-language models perform near chance on mental rotation tasks used in IQ tests, where humans typically score above 80%. The paper tested whether machines could identify the same 3D object from different angles—an ability architects and engineers rely on daily.

Classic logic puzzles also reveal flaws. In a 2024 experiment by Google and the University of Illinois Urbana-Champaign, models trained on variations of the Knights and Knaves riddle often repeated memorized solutions rather than adapting to new constraints. “Models depend on surface patterns rather than genuine reasoning,” the authors wrote in Nature Machine Intelligence.

Scale magnifies the problem. Apple researchers reported in February 2025 that large language models solve simple Towers of Hanoi and river-crossing puzzles flawlessly with small numbers of disks or travelers. Once complexity rises—six disks or six travelers—accuracy drops below 40%. A separate study by the University of Washington, Stanford University, and the Allen Institute for AI showed similar breakdowns on logic grid puzzles as the number of attributes and people increased.

For Malaysian businesses, the findings explain why AI cannot yet replace human intuition in areas such as design, diagnostics, or strategic planning where spatial manipulation and adaptive reasoning matter. While models may soon master trivia and text-based benchmarks, tasks requiring spatial ability or robust deductive logic remain beyond their reach.

Related: Malaysia Digital Economy Corporation (MDEC)

Reporting based on MIT Technology Review. Figures and claims are subject to revision as the story develops. DomainFork publishes editorial context, not investment advice — see our editorial standards.