Wafer Raises $40M to Automate AI Inference Optimization
Wafer just closed a $40M Series A (co-led by Marathon & Chemistry) to build AI that continuously optimizes AI inference workloads.
The wedge: Manual inference optimization is slow, expensive, and one-time. Wafer's autonomous agents learn traffic patterns and find optimal deployments across models, engines, kernels, and hardware.
How they won:
• Started with open-source LLM serving (YC S25)
• Proved 80% of Nvidia B200 performance on AMD at half the cost
• Landed customers trusting them with production workloads
• Attracted angels: Jeff Dean, Guillermo Rauch, Matthew Prince, Kyle Vogt
The metric: AMD MI355X now runs GLM5.2 at 2,626 tok/s/node—competitive with Blackwell.
Why it matters: Every dollar saved on inference scales across millions of requests. Founders: if you're running LLMs, this is your leverage.
https://www.wafer.ai/blog/series-a
#AIInfra #Founders
Post #404
182