Skip to content
← All news
3 min read

AMD is buying a startup that etches AI models directly into silicon

Taalas hardwires model weights into the chip instead of streaming them from memory. One chip runs one model, and runs it very fast.

The tradeoff is brutal and deliberate: the chip runs one model, nothing else, and it flies.

AMD announced on August 6, 2026 that it has signed a definitive agreement to acquire Taalas, a Toronto chip startup founded in 2023 by Ljubisa Bajic, who previously founded Tenstorrent and worked at both AMD and Nvidia. Terms were not disclosed. The team joins AMD's AI Group at close, subject to regulatory approval. Taalas had raised roughly $219 million: a $50 million seed when it came out of stealth in February 2024, and $169 million in February 2026.

What hardwiring a model actually means

A normal accelerator stores model weights in high bandwidth memory and streams them to compute units on every forward pass, which is why so much inference work is really memory bandwidth work. Taalas builds what it calls model-specific integrated circuits: the weights are etched into the silicon itself, so they never move. Its HC1 test chip, on TSMC's 6 nanometer process, is claimed by the company to serve Llama 3.1 8B at around 17,000 tokens per second, and a second chip targets models around 20 billion parameters. The tradeoff is inherent and not a bug: one chip serves one specific model.

Treat the speed multipliers as marketing

Coverage of the deal carries wildly different comparisons, one outlet reporting 48 times faster than Nvidia GPUs and another 73 times an H200 at a tenth of the power. Those are company benchmark claims measured against different baselines, and the gap between them is the tell. The checkable facts are the acquisition, the funding history, the process node, and AMD's stated plan: pair the technology with Instinct GPUs and Epyc CPUs inside its Helios rack systems in a disaggregated setup, GPUs handling prompt processing and hardwired chips handling token generation, programmed through ROCm.

Why a build studio cares

This is a bet that the model layer is stabilising. Etching weights into silicon only makes economic sense if a model is going to be worth serving, unchanged, for long enough to amortise a tape-out, which is a striking thing to wager in a market that has shipped three frontier releases in the last week alone. If that bet pays off, serving a popular open-weight model gets dramatically cheaper and the calculus of self-hosting shifts again. If it does not, it is expensive silicon for a model nobody runs any more. Either way the answer arrives on hardware timescales, not release-note ones.

Next step: read AMD's announcement and The Register's technical writeup. If you are weighing self-hosting against an API for a product, write to us at hello@gattyworks.com.

AMDAI ChipsInferenceAMDTaalasTenstorrentAIChipsInferenceSemiconductorsROCmHardwareAIInfrastructureAI

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.