Back to all posts

Etched Just Raised $300M at a $10.3B Valuation for a Chip That Can Only Run Transformers — And It's Beating Nvidia's Blackwell by 10x

Published on Jul 25, 20265 min read
GenAIDeveloper ToolsBackend

On July 23, 2026, Etched — a three-year-old chip startup that spent most of 2023 unable to get a single investor on the phone — closed a $300 million Series C led by Sequoia Capital, joined by Andreessen Horowitz, Jane Street, Diffusion, and SK Hynix. The round values the company at $10.3 billion, which Sequoia says is the highest valuation it has ever backed at the Series C stage. The bet behind that number is narrow and aggressive at once: Etched's Sohu processor runs exactly one kind of AI model — the transformer — and, according to the company's own benchmarks, does it dramatically faster than anything Nvidia currently ships.

A Round Sequoia Has Never Priced This High

Etched only came out of stealth on June 30, 2026, when it disclosed $800 million raised to date, a working Sohu chip, and more than $1 billion in signed customer contracts. Less than a month later, the new $300 million tranche pushes total funding past $1.1 billion and roughly doubles the company's valuation from the figure investors were reportedly discussing earlier in July. For a company that had not shipped a commercial chip a year ago, that trajectory says as much about investor appetite for an alternative to Nvidia as it does about Etched's own technology.

How Sohu Gets to 90% FLOP Utilization

General-purpose GPUs run transformer inference as software on programmable compute units, which is flexible but wastes most of the chip's theoretical throughput — Etched puts typical GPU FLOP utilization at 30-40%. Sohu takes the opposite approach: it hard-codes transformer attention directly into silicon as fixed-function logic instead of general-purpose compute, which the company says lets it use roughly 90% of the chip's FLOPS. Each Sohu chip carries 144GB of HBM3E memory — more capacity than an H100's 80GB — with about 1.8x the memory bandwidth of an H100 SXM5. The trade-off is architectural lock-in: a chip built to do one thing extremely well can't run anything that isn't a transformer.

The Benchmark Nvidia Has to Answer

In Etched's published comparisons, an eight-GPU cluster of Nvidia H100s serves Llama-3 70B at roughly 25,000 tokens per second, and an eight-GPU cluster of Nvidia's newer B200 'Blackwell' chips reaches about 43,000 tokens per second. An eight-chip Sohu cluster, Etched says, hits 500,000 tokens per second on the same model — roughly 20x the H100 cluster and around 10x Blackwell. Those numbers come from Etched's own benchmarking rather than independent third-party testing, which is worth keeping in mind, but the gap is large enough that it has drawn public responses from AI infrastructure analysts across the inference-hardware field, not just Nvidia's usual critics.

Two Harvard Dropouts Nobody Would Fund in 2023

Etched was founded in 2022 by Gavin Uberti, Chris Zhu, and Robert Wachen, who left Harvard's combined bachelor's/master's programs mid-degree and were named 2024 Thiel Fellows. By their own account, 2023 was close to a dead end: investors wouldn't take their calls, and the startup was running low on cash betting that transformers — rather than some future architecture — would remain the dominant way AI models are built. That bet is now the entire thesis behind a $10.3 billion valuation, and it's a specific kind of bet: Sohu only pays off for as long as transformers keep winning.

The Trade-Off: A Chip That Bets the Architecture Never Changes

The same specialization that gives Sohu its speed is also its biggest risk. Nvidia's GPUs are slower per chip on transformer workloads, but they run any model architecture a research team ships next — the entire reason GPUs became the default AI substrate in the first place. If a lab shipped a genuinely different dominant architecture, Sohu's fixed-function silicon would have nothing to fall back on, while a GPU fleet would just run the new workload less efficiently. Etched's counter-bet is that transformers are now too deeply embedded in production AI — and too economically important to displace on a hardware-refresh timeline — for that risk to matter within the life of this chip generation.

What It Means for Developers and AI Teams

For teams running LLM inference at scale, the relevant number isn't Etched's valuation — it's tokens per dollar per watt, and long-horizon agentic workloads that generate far more tokens per task than a single chat completion make that number matter more every quarter. If Sohu's real-world performance holds up outside Etched's own benchmarks, it changes the cost math for anyone serving high-volume transformer inference, particularly agent and coding-assistant workloads that already dominate token consumption. The near-term catch is access: Etched's disclosed contracts total in the billions but its customer list is still short, so most teams evaluating inference infrastructure today are choosing between GPUs, not between a GPU and a Sohu rack — that changes only as availability broadens beyond Etched's largest early customers.

Bottom Line

Etched just turned a company that couldn't get investor meetings in 2023 into a $10.3 billion bet that transformer-specific silicon beats general-purpose GPUs on the metric that determines inference cost at scale — FLOP utilization. The Sohu numbers, if they hold up under independent testing, describe a real shift in how AI inference economics could work; the risk sitting underneath them is that Etched has built a chip with no plan B if the industry's dominant architecture ever moves past the transformer. For now, that's a bet Sequoia was willing to price at the highest Series C valuation in its history.