A distributed compute framework for scaling AI workloads. Created and developed by @anyscalecompute.

Meet the Anyscale team at Fully Connected 2026 at Booth #600 and catch our session on Data Pipeline at Petabyte Scale: Production Video Curation with Ray on CoreWeave na2.hubs.ly/H07Z75-0
2
1
10
659
Ray Summit 2026, in two minutes. Three days in San Francisco. Keynotes on where Ray and Anyscale are headed. Breakouts from @Netflix, @Pinterest, @Spotify, @Zoox and more. The @vllm_project Conference. And the community that builds Ray. Every recording is now live on our YouTube channel: keynotes, breakouts, vLLM sessions, and Day 0 training. Watch on demand: na2.hubs.ly/H07TdkX0
3
1
15
1,302
Ray just works. If you grok the architecture and use Ray correctly, it scales very well. Thanks @robertnishihara and @anyscalecompute team!
Incredible work, and very cool to see @periodiclabs using @raydistributed.
6
38
6,329
Incredible work, and very cool to see @periodiclabs using @raydistributed.
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
1
4
37
10,991
Ray Summit 2026 recap is live. 2,000 attendees, three days, one theme: RL is a production infrastructure problem now. Lila Sciences, Torc, NVIDIA, Periodic Labs, Bedrock Robotics, Spotify, Capital One, Microsoft, Recursion, and more. Read the full recap. na2.hubs.ly/H07FqbB0
4
6
23
1,654
ray retweeted
How does one RL post-train a 397B model for long-horizon knowledge work? 👩‍💼 We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from Mercor Research on open model training research. Full blog: mercor.com/blog/training-fro… Source code: github.com/Mercor-Intelligen…
16
117
790
155,442
An agentic RL training guide for Qwen 3.5 397B with SkyRL. They show very large improvements with roughly 2000 expert-labeled tasks.
How does one RL post-train a 397B model for long-horizon knowledge work? 👩‍💼 We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from Mercor Research on open model training research. Full blog: mercor.com/blog/training-fro… Source code: github.com/Mercor-Intelligen…
6
3
42
5,652
That's a wrap on Ray Summit 2026. Three days, 1800+ attendees, 90+ sessions, Training Day, and the first vLLM Conference. Thank you to everyone who built it with us. Keynote replays are live: Day 1: na2.hubs.ly/H07yZF10 Day 2: na2.hubs.ly/H07yTZM0
3
5
19
1,451
ray retweeted
We stress tested Ray-2.58.0 for improvements on fault tolerance and reliability by scaling video data processing pipeline to caption 0.6 Petabytes of video data on 1600 Nvidia Blackwell GPUs in collaboration with @CoreWeave With the new release Ray Data shows: • 𝟲𝟬% higher throughput for streaming data processing • 𝟮𝟯% faster for batch inference • 𝟮𝟰% faster for Ray Data shuffle This will significantly improve how world model and physical AI companies handle pre-training data. Blog: coreweave.com/blog/video-cap…
5
17
1,023
ray retweeted
Proud to share a snippet of our ML progress in the Ray Summit keynote today by @LiamFedus. It’s been a lot of fun landing inference improvements and training an RL model that surpasses frontier models at a fraction of the cost. More details to share soon!
I got some as well. Inspirational talk! Also, very cool to see that @periodiclabs is using Kimi to reduce inference costs 20-50x over frontier APIs. Everyone will do this.
3
6
51
4,496
Ray passed 170M downloads in a single quarter, 5x growth in a year. The Ray Summit keynote covered why: teams are moving from using models to building their own, and closing the loop so they keep improving. Where AI infrastructure goes next: na2.hubs.ly/H07sTb30
3
5
18
1,341
Announcing Anyscale GPU Health Observability. GPU faults have always looked identical to broken scripts. Now they don't. DCGM signals correlated straight to the job and workspace running on that GPU, automatically, across KubeRay and VM. na2.hubs.ly/H07px5N0
3
4
19
1,626
Excited to share our work on large-scale sharded weight transfer! We’ve implemented a native sharded weight transfer engine in vLLM and SkyRL using Ray Direct Transfer (RDT) + NIXL. We’re able to achieve weight transfer for Kimi K2 (1T params) in BF16 in 7.53 seconds on 48 8xH100 nodes. Blog: vllm.ai/blog/2026-08-22-rdt-…
3
21
104
22,275
ray retweeted
Day one at @raydistributed and @vllm_project Summit ⚡ and we're on the schedule today 3 PM, Lightning Theater: Helen Zhao (@RedHat_AI) on scaling drafter training in real time across a full Verda GB300 rack with vLLM, Mooncake & Speculators. Booth's open all day on the expo floor. Come around, grab some Finnish chocolate 🍫 and talk shop: GPUs, distributed training, scaling AI in production.
1
4
9
1,325
Replying to @torc_robotics
From a laptop to a large production run, Torc uses the same stack to design their training. Rooted in Ray & @anyscalecompute
1
3
3
504
Really enjoyed this talk from @AndrewLBeam at @LilaSciences at Ray Summit today on the path toward automating scientific research: intelligence, software, & robotics.
3
4
23
4,134
Very happy to announce the release of Ray Sandboxes, developed with @a_sykim and his team at Google. It makes it easy to scale up your sandboxes with Ray, isolated by gVisor. It is fully open-source, check it out at anyscale.com/blog/announcing… Would love to get your feedback and contributions.
3
21
81
7,289
A higher KV cache hit rate doesn’t always mean lower TTFT and TPOT, or higher throughput. We studied LLM routers with different request scoring objectives under agentic and RL workloads. Here’s what we learned. 🧵
2
20
76
6,527