Meta’s AI chip roadmap runs through the cost line behind 8,000 job cuts
Arke is slated for data centers in the first half of 2027, followed by Astrid by year-end; Meta says the chips can cut total cost of ownership 44% on targeted workloads.
Meta Platforms plans to begin putting its third-generation in-house AI chip into data centers in the first half of 2027, a timeline that lands directly on a cost base Mark Zuckerberg has already described to employees as a zero-sum choice against payroll. The chip, MTIA 450, is codenamed Arke, Bloomberg reported. Its successor, MTIA 500 or Astrid, finishes design work in about a month and reaches data centers at the end of 2027, deployed more widely than anything in the line so far.
At an April town hall, Zuckerberg put the trade-off plainly. "We basically have two major cost centres in the company: compute infrastructure and people-oriented things," he said, according to Reuters. "If we're investing more in one area to serve our community, then that means we have less capital to allocate to the other. So that means we do need to take down the size of the company somewhat." Meta cut about 8,000 jobs beginning 20 May, roughly 10% of a workforce of 78,865, and cancelled 6,000 open requisitions.
The silicon is an attempt to bend the other side of that equation. Meta says the MTIA chips deliver a 44% reduction in total cost of ownership and a 40% improvement in power efficiency against general-purpose GPUs, for the specific workloads they target: ranking and recommendation, ad optimization and generative AI inference. Those are company figures for a narrow set of jobs, not a company-wide result.
Inference is the workload that keeps running after a model is trained, and Meta is trying to move more of that work onto hardware built for its own systems. Arke and Astrid are aimed at inference, while Meta still relies on Nvidia and AMD GPUs for training and broader AI infrastructure, and will keep buying GPUs from both in large volumes.
The deployment commitment is where the number becomes a plan. Meta has committed to deploying more than 1 gigawatt of its in-house chips within 12 months, after which the pace is expected to accelerate, a forecast Song said assumes no sharp downturn in AI demand. For scale, an internal memo reviewed by Reuters in July showed Meta deploying seven gigawatts of computing infrastructure this year and planning to double that to 14 gigawatts in 2027.
The spending has not slowed to meet the savings. Meta raised full-year 2026 capital expenditure guidance to $125 billion to $145 billion, up from $115 billion to $135 billion, sending shares down 9% after hours. To hold the build schedule it has locked in long-term supply agreements for memory with Samsung Electronics, flash storage with Sandisk and fiber-optic equipment with Sumitomo Electric, the July memo showed, as a memory shortage pushes component prices up across the sector.
How seriously Meta takes the cost arithmetic shows up in what it killed. A chip codenamed Olympus, designed to handle both training and inference, was cancelled so the program could concentrate on inference. Yee Jiun Song, Meta's vice president of engineering, said that at multi-gigawatt scale a dual-purpose chip costing roughly 30% more would be "completely unacceptable," according to TradingKey's account of his interview remarks. Song also said each generation takes on more technical risk in exchange for better performance per watt and per dollar.
The program has moved faster than the industry norm, and there is a documented reason. Meta laid out a four-generation roadmap in March covering the MTIA 300, 400, 450 and 500, targeting a new generation roughly every six months against an industry cycle of one to two years, enabled by a modular chiplet design. Broadcom co-designs the chips and TSMC manufactures them. TSMC delivered a first batch of 12 of the new chips on 1 September; actual performance came within 2% to 3% of simulation, the engineering team ran Meta's own models on them the same day, and initial testing turned up no design flaws, though months of debugging and optimization remain and foundry yield is still ramping.
The earlier generation shows why custom hardware buys something GPUs do not. MTIA 300, described in a post on Meta's engineering blog, puts network interface chiplets inside the chip package and hands collective communication to 16 dedicated message engines rather than the compute grid. Meta's engineers wrote that running large matrix operations alongside collectives degrades compute throughput by less than 0.5%, against more than 20% on traditional GPUs. On a 150-billion-parameter production recommendation model across 40 accelerators, they reported communication time 3.9 times faster than the equivalent GPU cluster.
Meta also tried to take cost out of headcount more directly. Reuters reported in August, from internal documents and interviews with more than 20 people, that a project codenamed Project OT explored cutting some team sizes by as much as 60% and replacing the work with AI agents. Zuckerberg called off the planned second wave on the night of 19 May, hours before the first cuts went out. Meta acknowledged the project, said the 60% figure covered scenarios for certain teams only, and said leaders cancelled the second wave before deciding how many jobs it would have cost.
Zuckerberg has promised no further company-wide layoffs this year. Chief financial officer Susan Li told analysts Meta does not yet know what its optimal long-term workforce size is.
Sources (12)
Related coverage
The Morning Brief is coming soon
The HEADCOUNT
Morning Brief.
Get on the list for HEADCOUNT’s weekday briefing on the business of work.
- Weekday mornings, built from published HEADCOUNT reporting
- Evidence-backed — every item traces to sourced coverage
- The Signal: what the day's developments indicate for hiring