BOOM! MASSIVE AI POWER ON YOUR DESKTOP OR POCKET.
It will make doomers freak out.
AMD bought the company that etches entire AI models into silicon. This is not another GPU. This is a fundamental rewrite of how inference works — and it just punched a hole through the memory wall that every centralized AI lab has been hiding behind.
On August 6, 2026, AMD announced a definitive agreement to acquire Toronto-based Taalas. Official announcement:
newsroom.amd.com/news/amd-ac…
Taalas does not load model weights from HBM. It is the model. Weights and dataflow are burned into the transistors themselves. The first chip, HC1, runs Meta’s Llama 3.1 8B at roughly 17,000 tokens per second per user on a single TSMC 6nm die drawing about 200–250 watts. That is an order of magnitude faster than leading GPUs on the same workload, at a fraction of the power and cost (vendor figures).
17,000 tokens per second is faster than any data center.
A PCIe card. In a regular server. Or, soon, on your desk or in your pocket.
A data center in your pocket.
This is the kind of raw, localized horsepower that makes centralized control look like a 2023 strategy. When inference no longer requires a warehouse of HBM-stacked accelerators and a permission slip from a frontier lab, the gatekeepers lose their choke point.
HOW IT WORKS — HARDCORE MODELS
1. The model becomes the chip
- Trained weights are encoded directly into mask-ROM (via-ROM) in the top metal/via layers.
- Only two of ~100 layers change between models.
- Automated foundry flow: new model → custom silicon in ~2 months.
2. Dual-domain architecture
- Mask-ROM recall fabric: immutable base weights. Fully digital.
Claimed density trick: store 4 bits + perform the associated multiply in a single transistor.
- SRAM recall fabric: KV cache + LoRA adapters for limited fine-tuning and context.
- No HBM. No CoWoS. No liquid cooling required for the first-gen card.
3. Hardware specs (HC1)
- Process: TSMC N6
- Die: 815 mm², 53 billion transistors
- Power: ~200–250 W
- Form factor: PCIe card
- First model: Llama 3.1 8B (aggressive mixed 3/6-bit quantization)
- Next: HC2 targeting ~20B parameters; larger models via multi-card pipeline parallelism
4. The speed source
- Weights never leave the die.
- No memory-fetch cycle per token.
- Compute is hardwired to the exact dataflow of that one model.
- Result: sub-millisecond first-token latency and 10k–17k tok/s class throughput on a single card.
Trade-off: the finished chip runs only the model it was built for. New architecture = new tape-out. Taalas argues that is still cheaper and faster than training the next model from scratch.
This architecture does something the safety-and-control crowd did not price in: it makes high-quality inference cheap, fast, and physically distributed. You no longer need to phone home to a handful of companies that own the GPUs, the evals, and the usage policies. A popular open model can be turned into silicon and dropped into ordinary machines.
Agentic loops, local coding, writing, and analysis at speeds that feel instantaneous become possible without a data-center bill or an API key that can be revoked.
AMD plans to fold the technology into its Instinct + Helios + ROCm stack, pairing the specialized silicon with GPUs for hybrid systems. That is the enterprise path. The more explosive path is the one that leaks out the side: commodity cards running fixed, high-demand models at speeds previously reserved for the biggest clouds.
Centralized AI control assumed inference would stay expensive, centralized, and observable. Hardwired silicon breaks that assumption. The model is no longer software running on someone else’s computer. The model is the computer — and it can sit on your motherboard.