λ
ai
ai.lmbda.com
λ
ai • POST
OpenAI’s Jalapeño Chip Won’t Automatically Bring Sora Back — It May Make Codex Even Harder to Beat
OpenAI’s Jalapeño chip makes AI inference faster and more efficient, but its biggest beneficiary may be Codex rather than a revived Sora video product.
2026-08-30
Home / AI & Search / Post
OpenAI’s Jalapeño Chip Won’t Automatically Bring Sora Back — It May Make Codex Even Harder to Beat

OpenAI’s decision to discontinue Sora has produced an appealing theory: if expensive video generation lost an internal battle for scarce compute, perhaps the company’s new custom AI chip can eventually make video cheap enough to return. The timing makes the idea tempting. Sora’s web and app experiences ended on April 26, 2026, while OpenAI is now publishing the first measured results from Jalapeño, its custom inference accelerator developed with Broadcom.

But the most important consequence of cheaper inference may be the opposite of a Sora revival. Jalapeño was designed around large language models and interactive agents, precisely the workloads behind ChatGPT and Codex. If the chip makes those products faster and more efficient, it could raise the economic return on the strategy that displaced Sora in the first place. The real question is therefore not whether OpenAI can create more compute, but which products will capture the new capacity.

Sora’s shutdown exposed compute as a product decision

OpenAI’s own support documentation confirms that the Sora consumer experiences were discontinued on April 26 and says the Sora API will follow on September 24, 2026. The company also tells customers that purchased ChatGPT/Sora credits can be used for Codex. OpenAI’s discontinuation notice makes the product transition explicit even though it does not itself provide a detailed economic postmortem.

Reporting around the shutdown described a much harsher underlying calculation. TechCrunch, citing a Wall Street Journal investigation, reported that Sora had been costing roughly $1 million per day during part of its operation while usage had fallen sharply from its peak. Recent reporting has also described OpenAI as concentrating resources on enterprise products and coding as competition with Anthropic intensifies. Those figures should not be treated as a permanent cost structure for AI video, but they illustrate why a visually impressive consumer product can still lose an internal allocation contest.

Video generation is unusually demanding because the provider performs substantial computation before knowing whether a user will value the result. A creator may generate several clips, discard most of them and keep only one. Coding agents have a different value proposition: when an agent completes a software task, fixes a bug or saves an engineer hours of work, the customer can connect the AI’s cost to labor and business output much more directly.

Jalapeño is real, but it was built for language-model inference

OpenAI and Broadcom announced Jalapeño in June as the first accelerator in a multigeneration compute platform. OpenAI says the chip was designed from the ground up for current and future LLM inference, with the broader system spanning hardware, models, kernels, networking, memory and serving software. That distinction matters when discussing video: Jalapeño is not a publicly announced Sora accelerator.

On August 25, OpenAI published its first benchmark results. The company tested Jalapeño on the public InferenceX benchmark with GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, saying it achieved a better combination of throughput, power efficiency and latency across the tested operating range. On Kimi, OpenAI reported about 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. The company also says Jalapeño is rated at 700 watts while measured sustained power remained at or below 550 watts on the workloads tested.

These are company-reported results rather than proof that every OpenAI workload will see the same improvement. More importantly, the published benchmarks are language-model workloads. OpenAI has not said that Jalapeño is optimized for the architectures, memory patterns and media pipelines required by a future Sora-class video generator. A custom chip can lower parts of OpenAI’s inference bill without making the economics of high-quality video generation suddenly trivial.

The strongest beneficiary may be Codex

Jalapeño’s stated design goals align particularly well with agents. OpenAI emphasizes latency because an agent may execute many sequential model calls and tool actions; small delays can accumulate across a long task. In an essay accompanying the benchmark release, CFO Sarah Friar framed custom silicon as part of a feedback loop connecting data centers, chips, frontier models, products and rising demand.

That creates a counterintuitive outcome for anyone expecting efficiency gains to free a fixed pool of chips for abandoned consumer experiments. If Codex becomes cheaper and more responsive, OpenAI can allow agents to reason for longer, execute more steps or serve more developers. The value of allocating infrastructure to coding may rise at the same time that the cost per unit of inference falls.

This is a familiar version of the rebound effect often associated with Jevons paradox: improving the efficiency of a resource does not necessarily reduce total consumption because lower effective costs unlock additional demand. In AI, faster and cheaper inference can be absorbed by longer contexts, more agent actions, larger user populations, lower prices and entirely new products. Efficiency creates room, but successful software quickly learns how to occupy it.

A future video product would probably look different from Sora

None of this means consumer AI video is gone permanently. It means a comeback would need a product model that competes better for infrastructure. The most plausible version may be video as a metered capability inside ChatGPT rather than the resurrection of a standalone Sora social network.

ChatGPT already supplies identity, billing, conversation context and distribution. A creator could potentially develop a storyboard, generate reference images, revise prompts and render video inside one project instead of moving among separate applications. A credit system could also expose the physical cost of generation more clearly, with duration, resolution or quality consuming different amounts of capacity. OpenAI has not announced such a system for a future video product, so this remains a product hypothesis rather than a roadmap claim.

Integration would save organizational overhead as well as compute, but it would sacrifice something Sora offered: a dedicated creative community. A generation tool inside a private assistant is not the same as a feed where creators discover clips, techniques and each other. Rebuilding that social layer would require moderation, recommendations and community operations in addition to the underlying model.

Cheaper intelligence does not mean spare intelligence

The useful lesson from Jalapeño is broader than Sora. AI companies are moving from buying generic acceleration toward co-designing silicon, models and serving infrastructure because product strategy increasingly depends on the economics of each inference. OpenAI’s first custom chip gives it another lever over those economics, alongside model optimization, data-center expansion and pricing.

For video creators, that is reason for cautious optimism rather than evidence of an imminent Sora successor. OpenAI has announced no Sora 3, and its first custom accelerator is explicitly centered on LLM inference. The chip may eventually contribute to an infrastructure stack capable of supporting more multimodal generation, but the immediate strategic fit is stronger with ChatGPT and Codex.

That is why Jalapeño could make a Sora return simultaneously more technically possible and less strategically urgent. If custom silicon makes every valuable agent interaction cheaper, the threshold for an expensive consumer video product to win compute may rise rather than fall. Sora’s future, if it has one, will depend not on unused chips appearing in OpenAI’s data centers but on video becoming valuable enough to compete for capacity that other products can also use.

Related
same category