Skip to content
Breaking
DND seeks more bases and storage facilities to support AFP modernizationGross domestic product: Detailed results on economic performance in Q2 2026Construction orders rise 6.0% in JuneNutex Health says unauthorised party stole data from serversMalaysian firms warned over rising identity fraud in account creation and recoveryWhatsApp adds multiple passkeys for phishing-resistant sign-ins on iOS and AndroidMarimo patches high-severity flaw allowing pre-cell execution of MCP commandsApple unveils new Mac mini, Mac Studio with major chip upgradesHong Kong’s big trade deficit almost disappears as exports surgeApple unveils M5 Ultra and M6 chips for new Mac Mini and Mac Studio modelsOpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks showTransport Ministry urges public to use licensed, roadworthy vehiclesMalaysia logs unhealthy air quality in 30 areas as haze deepensIndonesian labour ministry and Japan’s JICA deepen safeguards for workers in JapanSellers on Shopee and TikTok call out auto-generated AI ads that mislead customersLay Hong Starts FY27 Steadily On Higher Revenue, RM10 Million ProfitAI companion robots gain ground in Malaysian homes as loneliness gap widensThe safety penalty: Reclaiming operational sovereignty in the age of AI50% tariff threat could curb Canada's access to US marketAI-enabled malware still largely a lab threat despite 405 samplesDND seeks more bases and storage facilities to support AFP modernizationGross domestic product: Detailed results on economic performance in Q2 2026Construction orders rise 6.0% in JuneNutex Health says unauthorised party stole data from serversMalaysian firms warned over rising identity fraud in account creation and recoveryWhatsApp adds multiple passkeys for phishing-resistant sign-ins on iOS and AndroidMarimo patches high-severity flaw allowing pre-cell execution of MCP commandsApple unveils new Mac mini, Mac Studio with major chip upgradesHong Kong’s big trade deficit almost disappears as exports surgeApple unveils M5 Ultra and M6 chips for new Mac Mini and Mac Studio modelsOpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks showTransport Ministry urges public to use licensed, roadworthy vehiclesMalaysia logs unhealthy air quality in 30 areas as haze deepensIndonesian labour ministry and Japan’s JICA deepen safeguards for workers in JapanSellers on Shopee and TikTok call out auto-generated AI ads that mislead customersLay Hong Starts FY27 Steadily On Higher Revenue, RM10 Million ProfitAI companion robots gain ground in Malaysian homes as loneliness gap widensThe safety penalty: Reclaiming operational sovereignty in the age of AI50% tariff threat could curb Canada's access to US marketAI-enabled malware still largely a lab threat despite 405 samples
AI Edge

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.

Source: TechCrunch · August 25, 2026 at 2:30 PM · AI-assisted report

Single-source
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
Photo: Wikimedia Commons — Axiata

KUALA LUMPUR, 25 AUGUST 2026 —

Listen to this article

DomainFork Audio · read aloud

Share

OpenAI unveils Jalapeño, a new inference chip that outperforms Nvidia Blackwell on Semianalysis benchmark

Market Impact

OpenAI announced its latest inference processor, Jalapeño, at the Hot Chips conference on Tuesday. In a press call, the company’s head of hardware, Richard Ho, said the chip delivers a “very, very significant performance advance” over the current state‑of‑the‑art, citing higher tokens per user and greater throughput per kilowatt on Semianalysis’s InferenceX benchmark. The benchmark comparison was against an Nvidia Blackwell system, the most advanced inference platform available today.

Jalapeño was first announced in October 2025 and was developed in close collaboration with Broadcom, with OpenAI’s own models assisting in the design. The company plans to roll the chip out in a multigenerational platform that will integrate AI models, chips and memory in a full‑stack approach. According to a blog post, the design focuses on reducing data movement and communication delays during the prefill and communication phases of inference, which are often bottlenecks.

“We designed Jalapeño to minimise data movement and communication delays,” the post said, noting that model state, including the KV cache, can be explicitly placed and kept local while the system activates the right combination of compute, memory and networking for each inference phase.

Ho said Jalapeño would be available in “very small volumes” by the end of 2026, with larger deployments expected in 2027. The company cautions that the competitive landscape may evolve by that time, as Nvidia and other vendors are likely to release newer chips. Nevertheless, the current benchmark results suggest that Jalapeño can serve more AI work per unit of power while also delivering lower latency responses.

Impact on the Malaysian market Malaysia’s AI ecosystem is still in its early stages, with most inference workloads handled by cloud providers that rely on Nvidia GPUs and other commercial processors. The introduction of a more power‑efficient chip such as Jalapeño could influence local data‑center operators to reconsider their hardware mix, especially for high‑volume, low‑latency applications such as real‑time translation, customer service bots and autonomous vehicle control.

However, the chip’s limited initial deployment and the need for specialized software support may delay widespread adoption in the country. Local semiconductor companies, such as Axiata’s data‑center arm and the newly formed Malaysia AI Innovation Hub, may monitor the performance metrics closely to assess whether Jalapeño’s architecture aligns with their infrastructure requirements.

Sector and company specifics OpenAI’s collaboration with Broadcom highlights a trend of AI firms partnering with established semiconductor manufacturers to accelerate hardware development. Broadcom’s expertise in high‑performance networking and memory interfaces complements OpenAI’s model‑centric design philosophy. The chip’s emphasis on reducing communication overhead aligns with the broader industry shift toward edge inference, where power and latency constraints are paramount.

While the benchmark results are promising, details on the chip’s silicon area, power envelope and cost per unit are not yet available, making it difficult to quantify the economic impact for Malaysian enterprises.

Outlook The next few years will see a rapid evolution in inference hardware. OpenAI’s Jalapeño, with its focus on power efficiency and low latency, could set a new benchmark for AI service providers worldwide. For Malaysia, the key will be to evaluate whether the chip’s architecture can be integrated into existing data‑center infrastructures and whether the cost‑benefit ratio justifies a shift from Nvidia‑based solutions.

As the chip moves from prototype to production, more detailed specifications and pricing information will be required to assess its suitability for the Malaysian market. Details not yet available.

Related: Axiata

Reporting based on TechCrunch. Figures and claims are subject to revision as the story develops. DomainFork publishes editorial context, not investment advice — see our editorial standards.