Scale-up possibilities for everyone. API Platform: platform.stepfun.ai/ HuggingFace: huggingface.co/stepfun-ai

Pinned Tweet
Introducing Step 5 Preview: Advancing the Pareto Frontier. Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. - 600B total / 27B active MoE, with 1M context + Vision - Substantially lower task cost at comparable intelligence - Broad software engineering capabilities with sustained execution over long horizons Try Step 5 Preview: platform.stepfun.ai Model page: stepfun.com/step-5-preview Open weights on Oct 15.
200
280
1,878
433,479
We’ve open-sourced onPanda 🐼 — the tool we use internally for LLM data annotation and model inspection. The workflow is simple: find an error, correct the token, and let the model continue. ✍️ Data annotation - 52% lower median annotation time vs. manual post-editing - SFT + preference data in one workflow, with high on-policy fidelity (ΔPPL <1% vs. the model’s resampling baseline) - Precise token-level supervision with paired positive/negative examples, plus agent-trajectory annotation across image, audio, and video 🔎 Model inspection and debugging - Inspect token probabilities and top-k alternatives, steer decoding token by token, and explore SVG generation, web development, and agent tasks directly in the browser. Try it (mobile-friendly): onpanda.diyer22.com Paper: huggingface.co/papers/2609.2…
I spent two years building this interactive tool to let you steer LLMs and agents at the token level. Introducing onPanda — a web app for token visualization & control, model inspection, data annotation, and more. Try it online (works on mobile): onpanda.diyer22.com/
12
37
330
17,505
Huge thanks to @diyerxx and the team for this work 🫡
1
1
15
960
Introducing Step Code v0.1.0. Swift execution. High token efficiency. Long-horizon reliability. Now open source under the MIT License. Step Code handles the full development loop—from reading and editing code to running tests and shipping—from one CLI. - 80.9% on Terminal-Bench 2.1 in evaluation - 73.3% on Multi-Frame, our 150-task long-horizon benchmark - One-command static site publishing with StepPage GitHub: github.com/stepfun-ai/Step-C…
30
44
453
44,401
Step Code also includes StepPage: publish a local static site to a shareable URL with one command. Version management and rollback are built in, so development, debugging, and delivery can stay in the same terminal.
1
1
18
1,588
Install Step Code on macOS, Linux, or WSL: curl -fsSL static-openapi.stepfun.com/s… | bash Windows: WSL is recommended. PowerShell support is currently in beta. Issues and pull requests are welcome: github.com/stepfun-ai/Step-C…
1
3
18
1,295
Step Code also includes StepPage: publish a local static site to a shareable URL with one command. Version management and rollback are built in, so development, debugging, and delivery can stay in the same terminal.
1
1
242
Install Step Code on macOS, Linux, or WSL: curl -fsSL static-openapi.stepfun.com/s… | bash Windows: WSL is recommended. PowerShell support is currently in beta. Issues and pull requests are welcome: github.com/stepfun-ai/Step-C…
1
247
Step 5 Preview pushes our intelligence–cost Pareto frontier outward. It comes in at 44 on the Artificial Analysis Intelligence Index, $0.71 per task. Thanks @ArtificialAnlys for putting Step 5 Preview through the full evaluation. More to come on Oct 15.
StepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations Step 5 Preview is @StepFun_ai's new flagship model, with 600B total and 27B active parameters, succeeding Step 3.7 Flash (released May 2026). It scores 44 on the Intelligence Index, level with Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key takeaways: ➤ Step 5 Preview costs ~2.8x less per Intelligence Index task than models at the same score. It costs ~$0.72 per task, against ~$2.00 for Kimi K3 (max) at the same score of 44 and ~$2.01 for GLM-5.3 (max) at 45. This is driven by pricing: at $1/$2.70 per 1M input/output tokens, it is priced below both on input and output. MiMo-V2.6-Pro is the one model that scores higher (46) at a lower cost per task ($0.13) ➤ Frontier reasoning is the standout strength, and where the jump from Step 3.7 Flash is largest. Step 5 Preview scores 46% on Humanity's Last Exam, in line with Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both are up sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Higher AA-Omniscience accuracy than GLM-5.3 at fewer parameters, but with more hallucination. At 600B total parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark measuring factual recall and hallucination, ahead of GLM-5.3 (max, 34%, 753B) and behind Kimi K3 (max, 48%, 2.8T). It attempts more questions than GLM-5.3 (68% vs 55%) and hallucinates more often when it does (43% vs 30%), landing at 16 on the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agentic evaluations are where Step 5 Preview lags peers at a similar Intelligence Index score. It scores 1,566 Elo on GDPval-AA, our primary evaluation for agentic performance, behind Qwen3.8 Max (1,668) and GLM-5.3 (max, 1,646). The gap holds on Terminal-Bench 4.0 (33% vs 39% and 42%), AA-Briefcase (1,432 Elo vs 1,640 and 1,525) and AutomationBench-AA (51% vs 56% and 62%) Key model details: ➤ Model Size: 600B total parameters, 27B active MoE model ➤ Context window: 1M tokens ➤ Multimodality: Text, image and video input, text output ➤ Pricing: $1/$2.70 per 1M input/output tokens, with cached input at $0.05/M ➤ Availability: StepFun first-party API, with open weights release planned for October 15th ➤ Licensing: Closed weights currently, with weights release planned for October 15th
16
8
184
10,883
Introducing Step 5 Preview: Advancing the Pareto Frontier. Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. - 600B total / 27B active MoE, with 1M context + Vision - Substantially lower task cost at comparable intelligence - Broad software engineering capabilities with sustained execution over long horizons Try Step 5 Preview: platform.stepfun.ai Model page: stepfun.com/step-5-preview Open weights on Oct 15.
200
280
1,878
433,479
Finance is a particular focus for Step 5 Preview. Financial work has to stand up to scrutiny. We evaluate Step 5 Preview on its ability to identify and verify reliable information, reconcile differences across reports, make assumptions explicit, and produce internally consistent forecasts and reproducible valuations. FinStepBench tests these capabilities across LiveSearch, CorporateValuation, and DeepResearch. We also evaluate Step 5 Preview on FrontierFinance across six investment use cases.
1
71
8,411
3
2
53
6,963
StepFun and ACE Studio present StepAudio 3 Music, StepFun’s first music generation foundation model — built to turn a prompt and lyrics into a complete song. Describe the sound you want: • Genre, mood and vocal character • Instruments, key and BPM • Song structure and arrangement Four workflows in one model: 🎤 Song generation 🎹 Instrumental generation 🔁 Music cover 🎙️ Vocal-to-song arrangement Powered by ABC-COT, it plans musical structure and arrangement before synthesis. Generate a version, rewrite the prompt, edit the ABC notation, and iterate toward the sound in your head. (ABC-COT API coming soon) Try the interactive demo: › static.stepfun.com/blog/step…
21
26
143
9,815
Introducing StepAudio 3, our new family of 5 audio models for real-time voice, speech recognition, speech generation, audio generation and music. Realtime ranks #1 on Artificial Analysis for both Conversational Dynamics (98.9%) and Speech Reasoning (99.7%). ASR reaches 1.7% WER, matching the best result on the leaderboard. Build voice agents that handle interruptions, reason while speaking, and call tools. Transcribe speech, generate expressive voices, and create full audio scenes and music. Available now: Voice AI Lab: audio.stepfun.ai/ Blog: static.stepfun.com/blog/step…
27
34
408
35,268
We’re bringing StepAudio 3 to San Francisco this Wednesday. Live demos, an open AMA with the StepAudio team, and a panel with leaders from PLAUD.AI, Cresta, Coval, SGLang-Omni & StepFun on what works, what still breaks in production, and where real-time AI interaction is heading next. 📍 Sep 16 · SF 🎁 $100 API credits + drinks & light dinner Join us ↓ luma.com/bkpe5h92
7
1
32
2,632