the new dawn is upon us, zero trust now applies to weights
How abliterated models can get you pwned 👾
We backdoored a 7B open model for less than $50, pointed Codex at it and it silently stole credentials the moment we used the trigger phrase. Success rate was 100% with zero false triggers on normal user prompts.
Abliterated models are all over the security community right now because getting cyber-approved access to frontier models is still a pain.
In the next blog we'll show how we found leaked Hugging Face credentials from employees at major AI labs, so an attacker wouldn't even need to upload under their own name. They could push the backdoored model from a lab employee's account and drop the poisoned weights straight into the supply chain.