Malaysian firms warned against rushing AI model choices for security operations
Cisco Talos Intelligence tested 66 model-and-reasoning combinations from Anthropic and OpenAI on a Unix log-review task that mimicked incident triage, but the “best” scorer cost $55 and took 32 minutes—far too slow for frequent SOC use.
Source: Cisco Talos Intelligence · August 26, 2026 at 11:58 AM · AI-assisted report
Single-source
KUALA LUMPUR, 26 AUGUST 2026 —
Cisco Talos Intelligence tested 66 model-and-reasoning combinations from Anthropic and OpenAI on a Unix log-review task that mimicked incident triage, but the “best” scorer cost $55 and took 32 minutes—far too slow for frequent SOC use.
Market Impact
The researchers found that higher reasoning effort rarely translated into better results: GPT-5.6 Luna’s score fell when effort rose, while Terra fluctuated and GPT-5.6 Sol peaked at 96 after multiple jumps. Anthropic’s Claude Opus 4.8 gained eight points from medium to high reasoning, then lost 9.5 points when pushed to xhigh.
Talos therefore rejected single-score rankings and built a four-variable Pareto frontier—score, cost, time and downside consistency—so organisations can pick the trade-off that fits their workflow instead of chasing the headline number.
“Choosing your model is not as straightforward as we had hoped,” Cisco Talos said in a technical report released on Thursday.
The corpus comprised 80,054 simulated log records in 20 formats, totaling 48 MB, generated by the open-source EvidenceForge tool and frozen at version 1.12.0. Each model ran five rounds in its native agent harness—Claude Code for Anthropic, Codex for OpenAI—then produced a synthetic-confidence score from 0 (real) to 100 (synthetic). A panel of four analyst personae—Threat Hunter, Network Forensics, Host/EDR and Detection Engineer—aggregated scores to a median condition result.
Scores varied by persona: Threat Hunter averaged 43, while Detection Engineer averaged 31. Within the same model and reasoning tier, Threat Hunter typically beat Detection Engineer by five points, though distributions overlapped and no role guaranteed the highest mark every time.
The frontier shows that the fastest frontier condition finished in 2 minutes at $0.90, while the highest frontier score reached 96 at $55 in 32 minutes. Most SOCs will need to set thresholds: discard any frontier condition that exceeds budget, latency or failure-rate limits, then choose the remaining model with the highest mean score.
“You cannot assume that a model’s performance scales with the reasoning level you use,” the report cautioned. “More effort means more cost but doesn’t always mean better results.”
For Malaysian businesses, the benchmark matters because SOCs and DFIR teams here increasingly automate triage with imported LLMs. A mis-chosen tier can inflate cloud bills while slowing incident response—a risk that rises as Malaysian firms adopt cloud-native security stacks.
Analysts said the results should prompt local CISOs to run their own small-scale trials before wide deployment.
“Malaysian enterprises should treat this as a call to benchmark rather than to buy the headline scorer,” said one Kuala Lumpur-based cybersecurity consultant who asked not to be named.
Related: Cisco Talos Intelligence · Kuala Lumpur