Skip to content
Breaking
AI models flub these intelligence tests. Can you fare any better?Malaysian firms warned against rushing AI model choices for security operationsMudslide from Nepal kills several at Gyirong Port trade hubUS Cold War aid turned Thailand into a regional growth hubApple’s new cleaning cloth now cheaper at RM39Malaysia’s 2026 GDP growth forecast raised to 5.4%Rupiah slips to Rp17,725 as Fed policy clues weighJakarta stocks plunge 1.48% as political transition and risks weighApple Mac Studio 2026 Malaysia: M5 Max and M5 Ultra, priced from RM10,999Sri Aman Highest, Here Are 17 IPU Unhealthy AreasSamsung’s Galaxy Z Series roadshow in Kuala Lumpur draws crowdsIsmail Sabri to face corruption charges tomorrowWoman released under parole gives birth in prison, vows to quit drugsGoogle To Invest US$2 Billion In Malaysian Data Centre Campus In SelangorBorneo Rising goes national with Sabah shoot, no KL audition neededBursa Malaysia stays higher at midday as regional markets climbIJM secures two high‑tech projects worth RM909.5 millionCourt dismisses Penang tycoon’s defamation appeal against state chief ministerSenator accuses Trump of making millions from Iran conflictThree more arrested over Linkedua road-closure race probeAI models flub these intelligence tests. Can you fare any better?Malaysian firms warned against rushing AI model choices for security operationsMudslide from Nepal kills several at Gyirong Port trade hubUS Cold War aid turned Thailand into a regional growth hubApple’s new cleaning cloth now cheaper at RM39Malaysia’s 2026 GDP growth forecast raised to 5.4%Rupiah slips to Rp17,725 as Fed policy clues weighJakarta stocks plunge 1.48% as political transition and risks weighApple Mac Studio 2026 Malaysia: M5 Max and M5 Ultra, priced from RM10,999Sri Aman Highest, Here Are 17 IPU Unhealthy AreasSamsung’s Galaxy Z Series roadshow in Kuala Lumpur draws crowdsIsmail Sabri to face corruption charges tomorrowWoman released under parole gives birth in prison, vows to quit drugsGoogle To Invest US$2 Billion In Malaysian Data Centre Campus In SelangorBorneo Rising goes national with Sabah shoot, no KL audition neededBursa Malaysia stays higher at midday as regional markets climbIJM secures two high‑tech projects worth RM909.5 millionCourt dismisses Penang tycoon’s defamation appeal against state chief ministerSenator accuses Trump of making millions from Iran conflictThree more arrested over Linkedua road-closure race probe
AI Edge

Malaysian firms warned against rushing AI model choices for security operations

Cisco Talos Intelligence tested 66 model-and-reasoning combinations from Anthropic and OpenAI on a Unix log-review task that mimicked incident triage, but the “best” scorer cost $55 and took 32 minutes—far too slow for frequent SOC use.

Source: Cisco Talos Intelligence · August 26, 2026 at 11:58 AM · AI-assisted report

Single-source
Malaysian firms warned against rushing AI model choices for security operations
Image: blog.talosintelligence.com

KUALA LUMPUR, 26 AUGUST 2026 —

Listen to this article

DomainFork Audio · read aloud

Share

Cisco Talos Intelligence tested 66 model-and-reasoning combinations from Anthropic and OpenAI on a Unix log-review task that mimicked incident triage, but the “best” scorer cost $55 and took 32 minutes—far too slow for frequent SOC use.

Market Impact

The researchers found that higher reasoning effort rarely translated into better results: GPT-5.6 Luna’s score fell when effort rose, while Terra fluctuated and GPT-5.6 Sol peaked at 96 after multiple jumps. Anthropic’s Claude Opus 4.8 gained eight points from medium to high reasoning, then lost 9.5 points when pushed to xhigh.

Talos therefore rejected single-score rankings and built a four-variable Pareto frontier—score, cost, time and downside consistency—so organisations can pick the trade-off that fits their workflow instead of chasing the headline number.

“Choosing your model is not as straightforward as we had hoped,” Cisco Talos said in a technical report released on Thursday.

The corpus comprised 80,054 simulated log records in 20 formats, totaling 48 MB, generated by the open-source EvidenceForge tool and frozen at version 1.12.0. Each model ran five rounds in its native agent harness—Claude Code for Anthropic, Codex for OpenAI—then produced a synthetic-confidence score from 0 (real) to 100 (synthetic). A panel of four analyst personae—Threat Hunter, Network Forensics, Host/EDR and Detection Engineer—aggregated scores to a median condition result.

Scores varied by persona: Threat Hunter averaged 43, while Detection Engineer averaged 31. Within the same model and reasoning tier, Threat Hunter typically beat Detection Engineer by five points, though distributions overlapped and no role guaranteed the highest mark every time.

The frontier shows that the fastest frontier condition finished in 2 minutes at $0.90, while the highest frontier score reached 96 at $55 in 32 minutes. Most SOCs will need to set thresholds: discard any frontier condition that exceeds budget, latency or failure-rate limits, then choose the remaining model with the highest mean score.

“You cannot assume that a model’s performance scales with the reasoning level you use,” the report cautioned. “More effort means more cost but doesn’t always mean better results.”

For Malaysian businesses, the benchmark matters because SOCs and DFIR teams here increasingly automate triage with imported LLMs. A mis-chosen tier can inflate cloud bills while slowing incident response—a risk that rises as Malaysian firms adopt cloud-native security stacks.

Analysts said the results should prompt local CISOs to run their own small-scale trials before wide deployment.

“Malaysian enterprises should treat this as a call to benchmark rather than to buy the headline scorer,” said one Kuala Lumpur-based cybersecurity consultant who asked not to be named.

Related: Cisco Talos Intelligence · Kuala Lumpur

Reporting based on Cisco Talos Intelligence. Figures and claims are subject to revision as the story develops. DomainFork publishes editorial context, not investment advice — see our editorial standards.