AI Lab developing unrestricted models & distributed inference ⟠ Over 5m monthly downloads on Hugging Face ⟠

Base
Training Dolphin on top of @arcee_ai’s Trinity Large Thinking 398B across 72× RTX 4090 48GB GPUs using Prime-RL
17
10
104
13,480
Dolphin Network - 2-week update 1.2 trillion tokens generated with Qwen 3.6 35B Peak of 126B tokens per day & 1.8m tokens per second 1046 GPUs online now w/ 49 TB of aggregate VRAM 612 H100s worth of idle GPU memory repurposed for inference datagen.dphn.ai
Dolphin Network V2 is live — our first major upgrade since launch The architecture was rebuilt from scratch in Golang -> - Auto-updates (no manual migrations) - NVFP4 as default on Blackwell GPUs - Higher utilization & network throughput via improved routing & load balancing
7
17
112
20,842
Will be running datagen on the network until we have sufficient capacity - at which point we will open up our inference API to the public In terms of what we are adding next -> - GLM 5.2 for 8xH200 / B200 / B300 nodes - Windows WSL + MacOS support - Inference API + Open Router - Audio gen - Image gen - AMD GPU + CPU nodes - Using idle network inference for RL rollouts during model training - Sharded inference with large models split between many consumer GPUs (see our initial report on sharded inference over the internet here drive.google.com/file/d/1bzn… )
6
7
58
17,497
Dolphin Network V2 is live — our first major upgrade since launch The architecture was rebuilt from scratch in Golang -> - Auto-updates (no manual migrations) - NVFP4 as default on Blackwell GPUs - Higher utilization & network throughput via improved routing & load balancing
32
38
212
110,664
Dolphin X1 Trinity Nano is now live on @huggingface Our smallest decensored model yet - 6B MoE with 1B active parameters trained using only online RL Huge thanks to @TargonCompute for providing an 8xB200 node, @PrimeIntellect for hosted RL, and @arcee_ai for the Trinity series
10
29
183
39,285
Thanks for mentioning $POD Protocol-owned Uniswap v4 1% liquidity pools are a great way for early-stage projects to fund development
11
1,023
Qwen 3.6 35B data generation on Dolphin Network 22.8 billion tokens generated 383 GPUs online right now 24.33 TB of aggregate vRAM Equivalent to over 300 H100s worth of idle GPU memory repurposed for inference datagen.dphn.ai
Node provider rollout has been going well Our pool of inference nodes running Qwen 3.6 35B have generated over 3.2B tokens so far Total inference bandwidth -> 9400 t/s 28x RTX 4090 12x RTX 5090 8x RTX PRO 6000 & many other cards API access coming soon 🐬
25
21
202
53,287
We have a new paper on model weight proofs that run on every request inside the inference engine releasing soon In v1, we were relying on logprobs to verify the correct model is loaded via the expected fingerprint Alongside slashable bonds for operators, i’d argue that was already very strong but it will be greatly improved in v2 which we are releasing within a month For v2, we don’t have to rerun inference on a validator GPU in the same way you have to with logprobs - instead we have proofs that can be validated very quickly on a CPU with less overhead Preview of benchmarks attached
1
2
202
Guide on how to run a node in our docs dphn.ai/docs/running-a-node 60gb vRAM required to run in FP8 with full context We recommend 1x RTX 6000 PRO or H100 / H200 / B200 on @TargonCompute Smaller models for idle consumer GPUs coming soon ~~ Watch the 35B datagen live datagen.dphn.ai
5
1
18
6,565
Node provider rollout has been going well Our pool of inference nodes running Qwen 3.6 35B have generated over 3.2B tokens so far Total inference bandwidth -> 9400 t/s 28x RTX 4090 12x RTX 5090 8x RTX PRO 6000 & many other cards API access coming soon 🐬
Dolphin Inference Network node operation is now live for anyone who would like to beta test before we go into production $POD rewards live for testers Repurposing idle GPUs to run Qwen 3.5 35B MoE
39
18
177
150,303
Venice Uncensored 1.2 is live The model has been upgraded with vision support, a 4× larger context window & stronger tool-use capabilities Developed in collaboration with @AskVenice to deliver the most uncensored version of Mistral 3.2 24B Trained on Subnet 4 @TargonCompute
10
24
185
33,959
Install instructions 60gb vRAM required to run in FP8 with full context Suitable GPUs -> 1xRTX 6000 PRO 1xH100 / A100 80gb 4x3090 / 4090 2x5090 2xA6000 1xH100 / B200 / B300 We used 1xRTX 6000 PRO provided by @TargonCompute Full install instructions available in our docs dphn.ai/docs/running-a-node
2
2
10
3,911
Nodes will be processing synthetic data generation requests during the beta phase We have seeded it with 7m prompts from the CodeX-7M dataset The Qwen3.5 35B version of this dataset will be uploaded to HF upon completion Data generation stats can be viewed live at datagen.dphn.ai
1
1
11
3,919
Dolphin Inference Network node operation is now live for anyone who would like to beta test before we go into production $POD rewards live for testers Repurposing idle GPUs to run Qwen 3.5 35B MoE
8
10
69
78,168
Added remote monitoring of node software via our web app this week
2
1
10
3,695
Preview of our distributed inference software Community beta live in a few days Idle GPUs will be able to run models & earn $POD tokens for their contributions
11
12
77
15,149
Noticed you also had slow speeds on the Qwen-3.5-35B MoE That is probably related to ollama then Try llama.cpp or lmstudio
just a gentle reminder that nobody should use ollama > slower than llama.cpp on windows > slower than mlx on mac > slop useless wrapper > literal code thieves alternatives? > lmstudio > llama.cpp > exllamav2/v3 > vllm > sglang like literally anythingʼs better than ollama lmao
2
348
Read more about how we managed to train 405B with a single node here in our first blog post blog.dphn.ai/405b/
1
1
17
2,415
Dolphin X1 405B achieves SOTA on @NousResearch's RefusalBench, answering 100% of all commonly refused prompts Dolphin aims to put alignment in the hands of the user, without artificially restricting what prompts will or will not be answered by the model
1
1
14
2,355
Dolphin X1 405B is now live on @huggingface This model is a result of our efforts to decensor @allen_ai Tulu-3 405B efficiently, using just a single B200 Node to create the largest Dolphin model ever Thank you to @lium_io for the generous 8xB200 sponsorship
5
5
42
12,496
184 GPU nodes online generating 14m tokens per hour 24% through the dataset generation within 24 hours Contribute compute to earn $DPHN
Dolphin Network beta is now live We are building a distributed network of heterogeneous AI models running on machines all over the world - with verified inference to enforce trust Contribute compute & earn $DPHN token rewards on @Base Any GPU with ≥24gb of vRAM is eligible
12
7
51
6,962
How to run a node & earn $DPHN A step by step video tutorial is available - install -> generation time is around 10-15 minutes on a fresh Linux machine Instructions to run a node are located at dphn.ai/docs You can watch the data generation run live at datagen.dphn.ai
1
8
2,076
Why are we building a distributed inference network? Dolphin Network allows GPU owners to repurpose their idle compute by powering a portion of our inference network By processing the network's requests, nodes earn $DPHN tokens which can be used for inference or sold on the market
1
4
789
Nodes that contribute will be generating >1m compliant answers to commonly refused user prompts - which will help train future Dolphin models Data generation will stress-test the network and keep nodes under heavy utilization during the first stage of the rollout datagen.dphn.ai
1
5
714
Dolphin Network beta is now live We are building a distributed network of heterogeneous AI models running on machines all over the world - with verified inference to enforce trust Contribute compute & earn $DPHN token rewards on @Base Any GPU with ≥24gb of vRAM is eligible
22
14
113
23,723
Deep Infra offers excellent B200 access for $2.5 per GPU - @DeepInfra Amazing value offering and we are very grateful for their sponsorship
1
1
8
993
Dolphin X1 8B is now live on @huggingface This model is a result of our effort to directly uncensor Llama 3.1 8B instruct with our SFT + RL pipeline Special thanks to @DeepInfra for the generous 8xB200 sponsorship Try X1 8B for free in our Web chat UI or Telegram bot 🐬
10
12
124
12,215
Dolphin R1 Mistral 24B in at number #10 Try for free on @openrouter Graciously hosted by @chutes_ai
9 out of the 10 fastest-growing LLMs this week are open-source
3
33
5,980
Dolphin Mistral 24B Venice Edition v1.1 is now live The model has been retrained on an updated dataset w/ reinforcement learning used afterwards to improve creative writing capabilities & further remove censorship based on feedback from the previous version Developed in collaboration with @AskVenice this model delivers the most uncensored version of Mistral 24B for use across the Venice ecosystem Dolphin Mistral 24B Venice Edition is live as “Venice Uncensored” - the default model for all Venice users
7
17
115
8,281
URL scanning /scan [URL] is the format used to start a scan example = /scan bitcoin.org/bitcoin.pdf - Emulates Google Chrome to avoid rate limiting & captchas - Scanned text will stay in the model’s context memory & can be referenced in following prompts - We recommend using /clearcontext to wipe history when changing topics for higher quality responses
1
4
340
/menu is used to select between Logical & Creative mode + enable or disable context Disabling context means ever message is fresh and does not consider the conversation history above /clearcontext wipes the bots conversation memory which can be useful when changing topics The blue ≡ button in telegram lets you change between assistant & coding mode in the telegram menu
1
6
366
The Dolphin Telegram bot is now live Hosting Dolphin Mistral 24B Venice Edition for free to all Currently supports chat + coding configs & live URL scanning Usage guide below🐬🧵
2
16
73
7,011
Dolphin Mistral 24B Venice Edition is released Dolphin Mistral 24B Venice Edition is a collaborative project we undertook with @AskVenice with the goal of creating the most uncensored version of Mistral 24B for use within the Venice ecosystem Dolphin Mistral 24B Venice Edition is now live as “Venice Uncensored” - the new default model for all Venice users
23
33
184
21,960