Foundry Prices Jump 15% as AI Demand Tightens Capacity

  • Samsung raised 4nm, 5nm foundry prices up to 15% in July
  • Gartner: AI inference costs per agentic workflow rise fivefold by 2028
  • Marvell-Google TPU deal includes $12.2B stock warrant tied to chip purchases
  • Micron commits $10B over 10 years to new Boise research facility

Samsung Foundry’s 4nm production line in Pyeongtaek has been running at full capacity since late last year, producing logic chips for Qualcomm and base dies for high-bandwidth memory chips. In July, Samsung raised prices on new orders across its 4nm, 5nm, and 8nm foundry processes, with increases reaching 15% for customers in China and the U.S. U.S. export restrictions on chipmaking equipment have pushed Chinese firms to lean more heavily on foundries like Samsung, giving the company more pricing power in that market.

Inference costs jump even as token prices fall

AI inference costs per agentic workflow will increase more than fivefold through 2028, with the rate of innovation outpacing the cost curve. Gartner defines this as the “Inference Paradox” — better unit economics escalating the overall cost of AI without providing a clear pathway to commensurate and predictable value. While token costs will fall by 95% by 2030, inference costs for agentic workflows will increase more than fivefold over the next two years because AI app builders are using more, and often more expensive, tokens as large language models (LLMs) get ever more complex.

Gartner estimates that using an agentic reasoning model can increase provider inference costs by at least five times compared with a basic chatbot interaction, with more demanding workflows costing even more. This puts foundries building leading-edge AI chips in the rare position of raising prices mid-cycle. Typically, wafer pricing adjusts only with generational shifts. But with TSMC’s advanced nodes fully booked and demand overwhelming supply, Samsung and SMIC have both repriced upward. Samsung’s 4nm capacity is sold through 2027 at higher prices, and the company has been unable to meet all of the orders it receives.

Google diversifies TPU supply chain with $12.2B Marvell warrant

Marvell Technology has entered into an expanded commercial agreement with Google to develop custom semiconductor products across AI inference accelerators, storage controllers, network interface controllers, memory interface controllers and near-memory compute. In connection with the deal, Marvell has issued Google a warrant to purchase up to approximately 59 million shares at $206.58 per share — a potential stake valued at approximately $12.2 billion — with the warrant vesting primarily based on cumulative custom product revenue, one tranche for every $500 million in purchases by Google across 240 tranches through fiscal 2033.

The agreement positions Marvell as a second major custom silicon partner for Google alongside Broadcom, whose shares fell more than 3%. The scope covers not the TPU itself so much as the chips that move and store data around it. The deal adds Google’s Tensor Processing Unit to Marvell’s growing list of custom AI chip partnerships, alongside existing relationships with Amazon Web Services Trainium and Microsoft Maia.

Research facilities and advanced-packaging access expand

Micron plans to spend $10B over the next decade on a new research lab in Boise, Idaho, focusing on memory technologies, advanced memory and compute architectures, packaging, and manufacturing. The Semiconductor Research Corporation launched a $13M industry-driven project-based university research program spanning design, manufacturing, and packaging. UT Austin researchers are calling for broader access to advanced packaging facilities for universities, national labs and other researchers developing low-volume, high-value systems. The ‘Chips For Science’ report argues that expanded packaging infrastructure is needed to capitalize on U.S. semiconductor manufacturing investment, based on input from more than 70 experts.

The newly released National Security Science and Technology Strategy directs federal agencies to prioritize semiconductor R&D spanning electronic design automation (EDA), advanced manufacturing equipment and processes, heterogeneous integration and advanced packaging, AI/RF/optical hardware, 2D materials, and integrated photonics. Mitsubishi Chemical will commercialize ultra-high-purity colloidal silica for wafer polishing and chemical mechanical planarization (CMP) slurries used on interconnect layers, with new facilities at its Kyushu Plant expected to be operational by October 2028.

Key Takeaway

If your data-center capex assumes falling AI costs will offset rising workload complexity, rebuild those models. Foundries have repriced upward with capacity sold out through 2027, and Gartner’s fivefold inference-cost increase by 2028 means each new agentic workflow consumes far more compute than the simple chatbot it replaces. Plan for higher per-unit costs at the silicon level and escalating token consumption at the application level — both trends are moving against you simultaneously. The economics won’t improve until 2nm volume ramps or workload efficiency improves faster than capability growth, neither of which is visible in 2026 roadmaps.

Frequently Asked Questions

Why are foundries raising prices when token costs are falling?

Advanced-node foundry capacity is sold out through 2027, with AI chip demand outpacing supply. Samsung’s Pyeongtaek 4nm line runs at full capacity, and TSMC’s leading-edge nodes are fully booked. This gives foundries pricing power independent of downstream token economics. The 15% price increases apply to new orders only, avoiding friction with existing customers already locked into volume contracts.

What is the inference paradox and why does it matter for chip demand?

Gartner’s “Inference Paradox” describes how better unit economics escalate overall AI costs. While individual tokens become cheaper, agentic AI workflows use at least five times more computation than simple chatbots because they reason, replan, and call other agents continuously. This drives sustained demand for inference silicon even as model efficiency improves, keeping foundries at capacity despite falling per-token costs.


Article Source: Chip Industry Week In Review

Related posts