Dolphin Network - 2-week update
1.2 trillion tokens generated with Qwen 3.6 35B
Peak of 126B tokens per day & 1.8m tokens per second
1046 GPUs online now w/ 49 TB of aggregate VRAM
612 H100s worth of idle GPU memory repurposed for inference
datagen.dphn.ai
Dolphin Network V2 is live — our first major upgrade since launch
The architecture was rebuilt from scratch in Golang ->
- Auto-updates (no manual migrations)
- NVFP4 as default on Blackwell GPUs
- Higher utilization & network throughput via improved routing & load balancing
Dolphin Network V2 is live — our first major upgrade since launch
The architecture was rebuilt from scratch in Golang ->
- Auto-updates (no manual migrations)
- NVFP4 as default on Blackwell GPUs
- Higher utilization & network throughput via improved routing & load balancing
Will be running datagen on the network until we have sufficient capacity - at which point we will open up our inference API to the public
In terms of what we are adding next ->
- GLM 5.2 for 8xH200 / B200 / B300 nodes
- Windows WSL + MacOS support
- Inference API + Open Router
- Audio gen
- Image gen
- AMD GPU + CPU nodes
- Using idle network inference for RL rollouts during model training
- Sharded inference with large models split between many consumer GPUs
(see our initial report on sharded inference over the internet here drive.google.com/file/d/1bzn… )
🐬 Dolphin X1 Trinity Nano is HERE and it answers EVERYTHING
🔓 Built with a first-of-its-kind RL de-alignment pipeline — no hedging, no lectures, no dad advice
🔹 100% benchmark response rate vs GPT-5 at 11% and Gemini 2.5 Pro at 24%
🔹 Multi-gate, multi-judge reward system that blocks every escape route
🔹 Runs fully local on vLLM — your data never leaves your machine
🔹 Perfect for red teamers, security researchers and AI safety teams
🔥 Watch the full video below 👇
piped.video/-I31VXmicLk
Dolphin X1 Trinity Nano is now live on @huggingface
Our smallest decensored model yet - 6B MoE with 1B active parameters trained using only online RL
Huge thanks to @TargonCompute for providing an 8xB200 node, @PrimeIntellect for hosted RL, and @arcee_ai for the Trinity series
We have also released a blog post that goes into detail on the RL environment design, as well as the challenges we encountered along the way
blog.dphn.ai/dphn-x1-trinity…
You can also try it for free in our Web UI at chat.dphn.ai
Qwen 3.6 35B data generation on Dolphin Network
22.8 billion tokens generated
383 GPUs online right now
24.33 TB of aggregate vRAM
Equivalent to over 300 H100s worth of idle GPU memory repurposed for inference
datagen.dphn.ai
Node provider rollout has been going well
Our pool of inference nodes running Qwen 3.6 35B have generated over 3.2B tokens so far
Total inference bandwidth -> 9400 t/s
28x RTX 4090
12x RTX 5090
8x RTX PRO 6000
& many other cards
API access coming soon 🐬
Node provider rollout has been going well
Our pool of inference nodes running Qwen 3.6 35B have generated over 3.2B tokens so far
Total inference bandwidth -> 9400 t/s
28x RTX 4090
12x RTX 5090
8x RTX PRO 6000
& many other cards
API access coming soon 🐬
Venice Uncensored 1.2 is now live.
Developed with @dphnAI, this model delivers the most uncensored version of Mistral 24B.
Upgraded with vision support, a 4x larger context window, and stronger tool-use capabilities.
Trained on Bittensor Subnet 4 @TargonCompute.
First epoch of rewards for node providers has been paid out
50K $POD distributed to 33 providers based on relative contributions to the Qwen 35B inference pool
datagen.dphn.ai
Node provider rollout has been going well
Our pool of inference nodes running Qwen 3.6 35B have generated over 3.2B tokens so far
Total inference bandwidth -> 9400 t/s
28x RTX 4090
12x RTX 5090
8x RTX PRO 6000
& many other cards
API access coming soon 🐬
Node provider rollout has been going well
Our pool of inference nodes running Qwen 3.6 35B have generated over 3.2B tokens so far
Total inference bandwidth -> 9400 t/s
28x RTX 4090
12x RTX 5090
8x RTX PRO 6000
& many other cards
API access coming soon 🐬
Dolphin Inference Network node operation is now live for anyone who would like to beta test before we go into production
$POD rewards live for testers
Repurposing idle GPUs to run Qwen 3.5 35B MoE
Guide on how to run a node in our docs
dphn.ai/docs/running-a-node
60gb vRAM required to run in FP8 with full context
We recommend 1x RTX 6000 PRO or H100 / H200 / B200 on @TargonCompute
Smaller models for idle consumer GPUs coming soon ~~
Watch the 35B datagen live datagen.dphn.ai