The average GPU running in production sits at roughly 20% utilization — meaning companies are effectively paying for five times more compute than they use. Fixing that normally takes specialized engineers weeks of manual tuning, and the problem only gets harder as the number of model, engine and hardware combinations grows.

Wafer automates that optimization work with AI agents. Its autonomous agents learn a workload’s traffic patterns and performance constraints, then search for the optimal deployment across model, engine, kernel and hardware — what the company calls “AI that optimizes AI.” Wafer’s pitch centers on making non-Nvidia chips practical: in testing this past July, the company said AMD’s MI355X reached roughly 80% of the throughput of Nvidia’s B200 at less than half the cost on certain models, including GLM-5.2, and that its agents deliver 2 to 2.8x speedups over stock baselines across open-source models generally.
Wafer was founded by CEO Emilio Andere and co-founder Steven Arellano, both University of Chicago alumni. Arellano previously worked on Google Bard’s infrastructure and spent time at quantitative hedge fund Two Sigma. The pair raised a $4 million seed round in April, led by Fifty Years with participation from Liquid2 and Y Combinator, and angel backing from Jeff Dean and OpenAI’s Wojciech Zaremba.
Five months later, on September 1, Wafer announced a $40 million Series A co-led by Marathon Management Partners and Chemistry, with new participation from Wing Venture Capital, AMD Ventures and Outset Capital, alongside continued backing from Fifty Years and Y Combinator. The round values Wafer at more than $200 million — roughly a 50x jump from its seed valuation — and the company reportedly turned down acquisition offers from multiple cloud providers along the way.
The round also drew a notable roster of angel investors. Jeff Dean, who left Google after 27 years this past August to co-found Discovery Loop, joined alongside Guillermo Rauch (CEO, Vercel), Andy Fang (co-founder, DoorDash), Kyle Vogt (CEO, The Bot Company), Akshay Kothari (COO, Notion), Matthew Prince (CEO, Cloudflare) and Scott Stephenson (CEO, Deepgram).
Wafer said the new capital will go toward automating more of the inference optimization loop, so that every deployment gets the equivalent of a dedicated performance-engineering team continuously improving performance per dollar.
Open-source spinouts are driving inference-optimization competition
Wafer enters a field increasingly shaped by open-source projects turning into venture-backed companies. Inferact, built around the vLLM inference engine, raised a $150 million seed round earlier this year at an $800 million valuation. RadixArk, a commercial spinout of the UC Berkeley-born SGLang engine, was valued at roughly $400 million within months of launching. Baseten, a model deployment and serving platform, raised $300 million at a $5 billion valuation with Nvidia itself participating as an investor. Wafer differentiates itself by optimizing the deployment stack itself rather than shipping its own inference engine or serving platform.
MORE FROM THE POST
- Emerald AI Raises $150M Series A at $1.05B Valuation to Scale Power-Flexible Data Centers
- Groq Raises $350M Series A at $3.5B Valuation
- Groq Raises $650M to Pivot from Chips to AI Inference Cloud
- AI Security Firm HiddenLayer Raises $100M Series B
- Frontier Computing Raises $10M Pre-Seed to Grow AI on Living Neurons
Share
Most Read
- 1
- 2
- 3
- 4
- 5


Leave a Reply