Bittensor Subnet for AI Alignment | Operated by @astrowareai

Dubai
We’re excited to announce that Trishool’s HaloGuard 1.0 𝐡𝐚𝐬 𝐚𝐜𝐡𝐢𝐞𝐯𝐞𝐝 𝐒𝐎𝐓𝐀 prompt-safety performance among open-weight guard models. Today, we present HaloGuard 1.0, a constitutional input classifier for multilingual AI safety. It is built as a first-layer input guard that checks user prompts before they reach a downstream LLM, agent, or application. This is part of the safety infrastructure being built through @trishoolai , our decentralised AI red-teaming subnet on Bittensor SN23. Full arXiv paper goes live soon.
13
35
142
60,166
Most AI safety filters work like a bouncer with a list of banned words. See a bad word, block the message. Easy to fool. Ask one of these filters "how do I kill someone in Call of Duty" and it panics, because it sees the word kill and nothing else. Meanwhile a real threat, carefully worded to dodge any obvious trigger, walks right past it. That is the flaw in almost every guard model out there. It reads words. It does not understand what you mean. HaloGuard works differently. Instead of a banned-word list, it is trained on a constitution, a set of principles about what is actually harmful and why. So it does not scan for trigger words, it reasons about intent, the way a thoughtful human reviewer would. This is the approach Anthropic pioneered for Claude, and HaloGuard is one of the very few guard models built on that same foundation outside Anthropic itself. And attackers do not use banned words. They hide intent inside innocent-looking language, a birthday email, a roleplay, a harmless task. A word filter never sees it coming. A model that reasons about intent does. You cannot keyword-match your way to safety. Control has to understand, not just scan. That is what HaloGuard is built to do.
8
16
570
This month, three of the biggest names in security all moved on the same problem at once. That does not happen by accident. CrowdStrike shipped a product to secure AI agents at runtime. Proofpoint and Tenable each launched their own tools to inspect and control what agents are allowed to do. A startup built around AI red teaming got acquired. That is a category being born in real time. Agents did not fit the old security lanes. Endpoint, network, cloud, identity, none of them were built for something that reads, decides, and acts on its own across everything a company owns. So the industry is building a new lane for it, all at once, this month. Even CrowdStrike's CEO put it plainly. Governance alone cannot stop an agent already in motion. That is the whole point. You cannot write a policy and hope. You need a live layer that watches what the agent actually does and steps in before it acts. Control, not paperwork. We did not need this week to tell us that. We have been building that exact layer for months. Here is what the industry has not solved. Every one of those new products is owned by one company, a single vendor's closed box deciding what your agents can do. That is the same concentration of control we have been warning about all along. The category is real and it is here. The only question left is whether the layer that governs every agent should belong to one company, or to a network that no one owns. We already made our choice. Built on Bittensor, in the open, owned by no one. The market just confirmed the problem. We are building the answer you do not have to take on trust.
1
10
22
1,467
A grand experiment and an innovative approach to validation. Welcome inside the arena!
Subnet Summer started as a place to talk. Today it becomes a place where stake can speak. The Subnet Summer Validator is live, built with Medulla Labs. Where your root yield goes is your call. You vote, it moves. The Town Square of Bittensor now has a vote. 🔗 : subnetsummer.com/
1
1
15
914
Two weeks ago we shipped the first version of our output guard and showed you the starting number. 73.9% F1. We called it the floor, not the ceiling. Here is the floor already moving. The output guard is now at 78.53% F1. That is roughly four points from the current SOTA, and it climbed there the same way the input guard climbed to the top. The whole network attacking it, every week, every attack making it stronger. This is the part people underestimate about how we build. We do not ship a finished number and defend it. We ship a starting point and let the network drive it up in the open, where anyone can watch it happen. 73.9 to 78.53 in two weeks. The gap to the top is closing, and the loop that closes it does not slow down. The input guard already sits at the top of its class. The output guard is now knocking on the same door. Two guards. One at the summit. One climbing fast. Both built the same way, in the open, stronger every week. The control layer is not a promise about what we will build. It is something you can watch getting stronger in real time.
2
7
29
4,624
Last week we shipped the first version of our output guard and joined OpenAI's cyber program. Real progress on the control layer we are building. So it is worth saying why we build it as a network, not a company. Control is the layer that decides what every AI model is allowed to do. What it can say, what it can refuse, what it can act on. Whoever owns that layer owns the switch for the entire technology. So ask the obvious question. Who do you trust to hold that switch? Every answer that ends in one company is the wrong one. Not because those companies are evil, but because nobody should have that much power over something this important. The same labs asking the world to trust them are the ones whose models keep slipping loose. Control this important, cannot belong to anyone who can be bought, pressured, breached, or change their mind. It has to belong to no one. That is why it has to be a network. Owned by no company, running in the open, where no one can quietly flip the switch or weaken the guard. That is what Bittensor makes possible, and what Trishool is built on. The most important layer in AI cannot have an owner.
2
8
37
2,128
Read that email again. Nothing about it looks dangerous. A birthday message. A request to use someone's preferred name from the employee file. Polite, ordinary, the kind of thing that lands in a work inbox every single day. A human assistant reads it and thinks nothing of it. An AI agent reads it and does exactly what it says, and to finish the task it opens the employee file, reads personal records, and pulls out details it was never meant to expose. No hacking. No malware. No breaking in. Just a friendly message that quietly walks the agent into handing over data it should have protected. This is what makes agent attacks so dangerous. They do not look like attacks. They look like normal requests, and the agent has no instinct for when a task is a trap. That is exactly what HaloGuard 1.0 is built to catch. It sits over the agent and reads every instruction and every action before it happens, so a birthday email cannot quietly turn into a data leak. Your intern would pause. Your agent might not. Protect it.
1
9
32
1,645
Trishool | SN23 retweeted
Truly dedicated team in Bittensor.
We had a strong week. Accepted into OpenAI's Trusted Access for Cyber program, and the first version of our output guard trained and heading into the challenge. Any team moving this fast should get one honest question thrown at them. Are they actually in it for the long haul, or are they going to pump the story and quietly sell into it? You do not have to take our word for it. Check the chain. 125,000 alpha, locked perpetually. No decay, no unlock schedule, no end date. This is our own stake in SN23, locked for good. Not because anyone forced us to, but because you do not lock your alpha forever unless you intend to be here for everything that comes next. Anyone can talk. Locking your bag permanently, on-chain, for the whole world to see, is a different thing entirely. We are building the control layer for AI, and we are not going anywhere.
3
14
1,504
We had a strong week. Accepted into OpenAI's Trusted Access for Cyber program, and the first version of our output guard trained and heading into the challenge. Any team moving this fast should get one honest question thrown at them. Are they actually in it for the long haul, or are they going to pump the story and quietly sell into it? You do not have to take our word for it. Check the chain. 125,000 alpha, locked perpetually. No decay, no unlock schedule, no end date. This is our own stake in SN23, locked for good. Not because anyone forced us to, but because you do not lock your alpha forever unless you intend to be here for everything that comes next. Anyone can talk. Locking your bag permanently, on-chain, for the whole world to see, is a different thing entirely. We are building the control layer for AI, and we are not going anywhere.
𝟭𝟮𝟱,𝟬𝟬𝟬 𝗔𝗹𝗽𝗵𝗮. Locked perpetually. We just increased our conviction lock on SN23. We were already holding close to 100,000 Alpha under a decaying lock. We added another 25,000 and converted the entire position into a perpetual lock, with no decay, no unlock schedule, and no end date. 125,000 Alpha is now locked for good. We are doing this for one reason. We are building the control layer for AI, and we intend to be here for every part of it. While everything in Bittensor keeps shifting, we are doubling down with full commitment and no exit plan. On-chain proof below. 🔒: taostats.io/account/5CqTRwoF… 🔒: taostats.io/account/5Gyq1BBU…
3
9
59
4,430
For months we have talked about the output guard as the next layer. It is not a plan anymore. It is trained, and it goes into the challenge live. Here is what that means. The input guard controls what goes into the model. The output guard controls what comes out, every response the model generates and every action it proposes before it reaches anyone. It is the second wall, and it is a harder problem than the first, because judging what an AI produces is more subtle than screening what it receives. This is version one, and we are showing you exactly where it starts. F1 around 73.9%. Not SOTA yet, but that is the floor, not the ceiling. Remember how the input guard got to the top. It did not start there. It climbed, week after week, because the whole network attacked it and every attack made it stronger. The output guard now enters that same machine. The same challenge, the same global red team, the same loop that already produced one state of the art guard. That is the difference between a number and a trajectory. 73.9% is not where this ends. It is where the climb begins. Two guards now. One already at the top of its class. One just starting the same climb that took the first one there. The control layer is being built exactly the way we said it would be, layer by layer, in the open, and stronger every week.
4
7
31
1,252
At Trishool, we are currently focused on adoption, and we believe latency is a major part of adoption. We could have the most accurate guard model, but without the right response speed, it becomes unusable. That is why HaloGuard 1.0 is 0.8B, and why our streaming output guard is being built at the same scale. Most labs are building bigger guard models, like 4B, 7B, 8B, 12B, or even 27B parameters. We chose a smaller model, one-tenth the size of most competitors, that still beats them on benchmark F1. The idea is to compress larger-model safety performance into a smaller model that runs faster and cheaper. A sub-1B parameter guard can run on-device, on laptops, and on edge servers, not just in datacenters.
4
10
36
1,325
Trishool | SN23 retweeted
Trishool (SN23) has been accepted into OpenAI’s Trusted Access for Cyber programme. This is not a marketing badge. TAC is OpenAI’s vetted channel for providing more capable, less-restricted models to verified defenders working on real cybersecurity problems. @trishoolai is now part of that cohort. The more interesting part is what sits underneath the announcement. Trishool’s miners compete to break Halo, the subnet’s guard model. Successful attacks become training data for the next cycle, creating an incentive loop where the network continuously searches for new failure modes and improves the model. As frontier systems move into production, the need for serious security infrastructure will only increase. A Bittensor subnet being trusted by OpenAI to contribute to that layer is a meaningful commercial signal for both Trishool and the wider ecosystem. More to come from the team. Worth watching.
We are pleased to announce that Trishool (SN23) has been accepted into OpenAI's Trusted Access for Cyber program. This is a meaningful step in our commercial growth. Trusted Access for Cyber is OpenAI's vetted program that places advanced frontier capabilities in the hands of verified defenders doing serious cybersecurity work. Acceptance means Trishool now stands alongside some of the most respected names in security and enterprise as a trusted participant in the defensive AI ecosystem. As AI systems grow more capable and move deeper into production, the demand for a serious security layer grows with them. That is exactly where Halo, our guard model, comes in, and this program strengthens our ability to bring it to the organisations that need it most. This is also one part of a growing relationship with @OpenAI. We are actively exploring several directions together, with more to share in the months ahead. The work continues.
4
23
1,983
Trishool | SN23 retweeted
Just a reminder that Trishool’s HaloGuard 1.0 hit SOTA, outperforming models like ShieldGemma and NemoGuard. They were officially accepted into the Google for Startups Web3 Program. They were accepted into the Claude Partner Network. Recently accepted into OpenAI’s Trusted Access for Cyber Program. And they are now training a streaming output guard model that I believe has a real shot at reaching SOTA too. If this is not the definition of a bullish subnet on bittensor | $TAO, I don’t know what is.
We are pleased to announce that Trishool (SN23) has been accepted into OpenAI's Trusted Access for Cyber program. This is a meaningful step in our commercial growth. Trusted Access for Cyber is OpenAI's vetted program that places advanced frontier capabilities in the hands of verified defenders doing serious cybersecurity work. Acceptance means Trishool now stands alongside some of the most respected names in security and enterprise as a trusted participant in the defensive AI ecosystem. As AI systems grow more capable and move deeper into production, the demand for a serious security layer grows with them. That is exactly where Halo, our guard model, comes in, and this program strengthens our ability to bring it to the organisations that need it most. This is also one part of a growing relationship with @OpenAI. We are actively exploring several directions together, with more to share in the months ahead. The work continues.
1
6
16
1,278
We are pleased to announce that Trishool (SN23) has been accepted into OpenAI's Trusted Access for Cyber program. This is a meaningful step in our commercial growth. Trusted Access for Cyber is OpenAI's vetted program that places advanced frontier capabilities in the hands of verified defenders doing serious cybersecurity work. Acceptance means Trishool now stands alongside some of the most respected names in security and enterprise as a trusted participant in the defensive AI ecosystem. As AI systems grow more capable and move deeper into production, the demand for a serious security layer grows with them. That is exactly where Halo, our guard model, comes in, and this program strengthens our ability to bring it to the organisations that need it most. This is also one part of a growing relationship with @OpenAI. We are actively exploring several directions together, with more to share in the months ahead. The work continues.
10
19
106
12,015
Almost every major AI incident happened during model output or action, not at the prompt stage. Input guards cannot catch that because they read what a user asks for, not what the model does once it starts responding. That is exactly why we are building an output guard. Halo 1.0 is a pre-generation input guard, while Phase 3’s output guard is being built as a streaming model. With this output guard active alongside HaloGuard 1.0, we will be able to: ▪︎ Catch harmful responses mid-generation ▪︎ Stop dangerous tool calls before they execute ▪︎ Cut off sensitive data leakage the moment it starts Current Phase 3 miner challenges are producing the training data to get us there.
4
9
34
1,724
Trishool (SN23) is now successfully in the Arbos root basket. SN23 was initially scored out after our 80% burn adjustment following V440, which was misread as misaligned incentives. In reality, it was a deliberate mechanism update to protect the subnet below the demand bar. We shared the receipts, and Arbos revised after verifying that: ▪︎ The 80% burn is a real burn, not a reward ▪︎ Weights are updated weekly based on winners ▪︎ Validation code, miner scores, and weight settings are public and traceable We built Trishool to be transparent and permissionless, and we’re pleased to be included by the Arbos validator. 🔗: discord.com/channels/7996720…
4
5
40
1,284
Trishool | SN23 retweeted
HaloGuard cited as a reference point for model alignment safety alongside Meta's LlamaGuard and AI2's WildGuard. Trishool (SN23 on Bittensor), powers HaloGuard by providing the adversarial hardening layer to maximize true threat detection. bittensor:native
Quietly, something happened this week that matters more than any benchmark. HaloGuard showed up in someone else's research. A new paper on training safety guards, C-Guard, referenced HaloGuard as part of the foundation the field is building on. Not in passing, but in the table where researchers line up the guard models that define the space, next to work from Anthropic and the other serious names in AI safety. That is what a real contribution looks like. You do not get to decide you matter. Other people decide it, by building their own work on top of yours and measuring themselves against it. We did not ask for the citation. We did not know it was coming. A researcher we have never spoken to read the HaloGuard paper, found it worth referencing, and placed it alongside the constitutional work that started this field. This is how a control layer stops being a claim and becomes infrastructure. Not when we say it is important, but when the people building the next thing point back to ours. Still early. Still building. But the field is starting to notice. 🔗: arxiv.org/abs/2608.00180
1
7
64
5,288
All month we have been saying the control layer for AI cannot belong to one company holding the keys behind a black box. Easy thing to say, so let me show you we mean it. HaloGuard is open weights. The whole model is public, and anyone anywhere can download it, see exactly how it works, run it on their own machine, and try to break it. You do not need permission from anyone, and you do not have to take our word for anything. We did not do it for the marketing. We did it because a safety layer you are not allowed to look inside is just one more company telling you to trust them and leave it at that. That is the same posture as the labs whose models keep slipping loose while they insist everything is under control behind closed doors. If you are going to build the thing that decides what AI is and is not allowed to do, the whole world needs to be able to see it, test it, and poke holes in it. That is not a nice bonus on top. Without it there is no reason to believe the thing works at all. So go download it and go break it. That is the only way anyone should trust a control layer, including ours.
The best way to trust an AI safety model isn't to take our word for it. It's to test it yourself. That's why 𝗛𝗮𝗹𝗼𝗚𝘂𝗮𝗿𝗱 𝟭.𝟬 𝗶𝘀 𝗻𝗼𝘄 𝗽𝘂𝗯𝗹𝗶𝗰𝗹𝘆 𝗮𝘃𝗮𝗶𝗹𝗮𝗯𝗹𝗲 𝗼𝗻 𝗛𝘂𝗴𝗴𝗶𝗻𝗴𝗙𝗮𝗰𝗲 Any developer, researcher, or builder can download the model weights, evaluate them against their own workloads, and integrate HaloGuard directly into their stack without waiting for an API key or a commercial agreement. This wasn't an afterthought, it was a deliberate decision. We believe AI safety shouldn't be locked behind closed doors. The more people who can inspect, test, and build with HaloGuard, the stronger the ecosystem becomes. Open weights allow researchers to validate our work, builders to deploy it in production, and the community to uncover edge cases that make the model more resilient. That's how better safety infrastructure gets built. HaloGuard 1.0 covers 46 languages, is peer-reviewed, production-ready, and runs inline in under 100ms. Both the 0.8B and 4B variants are available, giving teams the flexibility to choose the model that best fits their deployment. Pull the model, run it against your stack, and tell us what you find. HuggingFace: huggingface.co/collections/a…
1
5
30
1,439
Trishool | SN23 retweeted
Are you still sleeping on Sn23?
Big News: One of Trishool’s top holders has locked approximately 2,000 TAO worth of SN23 alpha in Conviction. For a while, there were concerns around what could happen if a holder of that size exited without warning. This lock shows real belief in Trishool, and more importantly, it makes that alignment visible on-chain for everyone to verify. This is a call to the community to join us, participate, and support the long-term vision as we build Trishool into the next-generation safety layer for AI on Bittensor. This is the kind of conviction we want to see around SN23, because Trishool belongs on Bittensor and we are here to stay.
1
1
7
532
Big News: One of Trishool’s top holders has locked approximately 2,000 TAO worth of SN23 alpha in Conviction. For a while, there were concerns around what could happen if a holder of that size exited without warning. This lock shows real belief in Trishool, and more importantly, it makes that alignment visible on-chain for everyone to verify. This is a call to the community to join us, participate, and support the long-term vision as we build Trishool into the next-generation safety layer for AI on Bittensor. This is the kind of conviction we want to see around SN23, because Trishool belongs on Bittensor and we are here to stay.
5
9
43
2,738
Quietly, something happened this week that matters more than any benchmark. HaloGuard showed up in someone else's research. A new paper on training safety guards, C-Guard, referenced HaloGuard as part of the foundation the field is building on. Not in passing, but in the table where researchers line up the guard models that define the space, next to work from Anthropic and the other serious names in AI safety. That is what a real contribution looks like. You do not get to decide you matter. Other people decide it, by building their own work on top of yours and measuring themselves against it. We did not ask for the citation. We did not know it was coming. A researcher we have never spoken to read the HaloGuard paper, found it worth referencing, and placed it alongside the constitutional work that started this field. This is how a control layer stops being a claim and becomes infrastructure. Not when we say it is important, but when the people building the next thing point back to ours. Still early. Still building. But the field is starting to notice. 🔗: arxiv.org/abs/2608.00180
2
16
63
9,734