Technically is really really hard to stop distiliation even by stripping away the thinking tokens ? you could always use a SOTA open weight model and alongside humans it can generate the thinking tokens for that output , and you could automate that so you could extract significant data at scale ... maybe the SOTA AI models can go behind application layers instead ?
This explains a lot about why Anthropic hates Chinese models and open source.
Anthropic’s Nick Marwell says frontier AI could eventually require $10B, $100B, even $1 trillion training runs.
His argument: if a company spends that much pushing the frontier, competitors shouldn’t be able to distill the model and ship a cheap copy weeks later.
That’s the real fight.The problem is that not every Chinese model is just distillation. Chinese labs are spending serious money on training, building their own techniques, and getting extremely good at optimization and efficiency.
And as models keep improving, I think capabilities will increasingly converge. The gap between the top labs may become much smaller, which makes protecting one company’s frontier advantage much harder.