Gimlet Labs raised a $300 million Series B on September 4, 2026, led by Andreessen Horowitz. The San Francisco company said the money will expand a multi-silicon inference cloud that splits agent workloads across GPUs, near-memory chips, dataflow accelerators, and CPUs. New and returning backers named in the post include Sapphire Ventures, Menlo Ventures, 645 Ventures, Arm, Eclipse, Emergence, Factory, Hudson River Trading, M12, OnePrime Capital, Prosperity7, QuantumLight, Samsung Ventures, Tiger Global Management, Triatomic, Wing Ventures, and XTX Markets.
The raise comes a little more than five months after Gimlet’s Series A. Since March, the company said it has added billions of dollars in contracted revenue, a gigawatts-scale data center pipeline, and is “quickly scaling to hundreds of megawatts in managed capacity.”
Andreessen Horowitz published a matching note the same day. Growth partners Raghu Raghuram, Sarah Wang, Shangda Xu, and Stephenie Zhang wrote that Gimlet already counts a frontier lab and a hyperscaler as customers, and that the stack has delivered up to 10 times more throughput and interactivity on frontier models inside the same power envelope.
Why one chip type is no longer the pitch
Gimlet’s argument is that inference is no longer one job. Prefill is compute-heavy. Decode is memory-bound. Agents stack dozens of model calls, tool runs, and CPU work. The company said monthly token generation is already up 6 times in 12 months, and that models have grown from a few hundred billion parameters to 3 to 10 trillion, with context windows past 1 million tokens.
Homogeneous GPU clusters, it says, force a choice between high throughput and low latency. Gimlet traces a model, breaks it into phases, and schedules each slice on the silicon that fits the SLA. The blog lists prefill/decode splits, speculative-decode splits, and attention-FFN splits, and says the software can spin extra decode capacity on a second architecture when the first is full. The claimed result is 3 to 10 times faster frontier work, or 5 to 10 times more throughput for the same power.
a16z framed the same problem as a watt shortage. The firm said five U.S. hyperscalers are expected to spend $1 trillion in capex next year, and that adding plants and fabs will still lag software demand. It named Zain Asgar, Michelle Nguyen, Natalie Serrino, Omid Azizi, and James Bartlett as the team that already pushed from kernels and compilers into networking, power, cooling, and data center construction.
Decoded Take
This is not another GPU-cloud raise dressed up as research. Gimlet is selling a compiler-and-plumbing story: slice the model, mix the racks, and turn stranded watts into tokens. The $300 million check, plus a16z on the board, is a bet that the next bottleneck is orchestration, not another homogeneous hall. The risk is the same one every multi-vendor stack hits. Customers still buy NVIDIA by the acre, and mixing inlet temperatures, networks, and kernels is how outages start. Watch whether the unnamed frontier lab and hyperscaler show up in disclosed megawatts, whether Arm’s check turns into Arm silicon in the default pool, and whether those “billions” in contracted revenue convert before the next power interconnection slips.