The AI OS for scientific discovery. We run the whole research lifecycle — ideas, experiments, IP, funding, publication. Built for world-class science.

ScienceGuru retweeted
Here’s JEV-27B in action. Our open System 1 + System 2 model pairs calibrated, structured decisions with deliberate reasoning in one set of weights. Try the live demo on Hugging Face: huggingface.co/spaces/autotr…
Open source matters more when people can actually try it. Huge thanks to @huggingface and @multimodalart from the Hugging Face open-source team for building a live interactive JEV-9B demo on Spaces, powered by ZeroGPU. JEV-9B is our open System 1 model for calibrated, structured decisions. Try the demo: huggingface.co/spaces/autotr… huggingface.co/spaces/autotr… Model card: huggingface.co/autotrust/JEV… huggingface.co/autotrust/JEV…
2
7
211
ScienceGuru retweeted
Replying to @multimodalart
JEV-27B is now live to try on @huggingface. Give it a structured decision—yes/no, multiple choice, or a 0–5 rating—and see our open System 1 model return calibrated probabilities in one pass. Try the interactive demo: huggingface.co/spaces/autotr… Open weights + model card: huggingface.co/autotrust/JEV…
1
2
2
42
ScienceGuru retweeted
Open source matters more when people can actually try it. Huge thanks to @huggingface and @multimodalart from the Hugging Face open-source team for building a live interactive JEV-9B demo on Spaces, powered by ZeroGPU. JEV-9B is our open System 1 model for calibrated, structured decisions. Try the demo: huggingface.co/spaces/autotr… huggingface.co/spaces/autotr… Model card: huggingface.co/autotrust/JEV… huggingface.co/autotrust/JEV…
1
3
4
289
ScienceGuru retweeted
5/
Two systems, one model. System 1 answers typed questions in one pass with calibrated probabilities. System 2 is Qwen3.8-27B's reasoning, untouched: HumanEval 78.0%, byte-identical to the base. One engine, routed per request.
1
2
2
16
ScienceGuru retweeted
2/ Fidelity. On 25,376 held-out questions labelled with Jev 1.13's own output distributions, JEV-27B's mean KL divergence is ≈0.017. In plain terms: it takes ~60 sampled decisions to gather even one nat of evidence about which model answered.
1
3
2
69
ScienceGuru retweeted
1/ Jev showed that a model which decides, rather than writes, can move markets: its maker was reportedly in talks at a $10B+ valuation within 10 days of launch. Today we're open-sourcing JEV-27B. It reproduces Jev 1.13's decisions and adds full reasoning, on one GPU. 🧵
1
3
5
134
ScienceGuru retweeted
72.24 minutes to GPT-2 on 8×H100. AutoTrust's ScienceGuru, running Guru Turbo 1.2, posted the fastest result we've found on @karpathy 's Time-to-GPT-2 benchmark (one run, self-reported): 27% under the official record and 9.6 min ahead of the best community recipe. The field Reported times on the official leaderboard and in public nanochat PRs, as of Sept 24: - 72.24 min · ScienceGuru (AutoTrust) · 1 run - 81.84 min · Giovanni Zinzi · 6 runs - 91.74 min · Oriole Networks · 3 runs - 94.58 min · Martin Jurča (Seznam.cz) · 6 runs - 94.6 min · Weco-optimized run · 3 runs - ~99 min · Official record, Karpathy's autoresearch round 2 · 5 runs Zinzi also reported a 73.92-min experiment with a Cosmopedia data mix (3 runs), which he set aside after it regressed at a smaller model size. The lead - 27% faster than the official record (−26.8 min) - 12% faster than the best community recipe (−9.6 min) - 19–22 min ahead of Oriole Networks, Seznam.cz and the Weco-optimized run How Starting from nanochat and Zinzi's d22 recipe, @ScienceGuruAI narrowed the MLP (5,120 → 4,864) and set a 9,841-step training horizon. FP8, FlashAttention 3 and Muon/AdamW are unchanged. CORE: 0.259212, above GPT-2's 0.256525. Why it matters The official record came from Karpathy's own autoresearch agents. Ours was researched and coded by ScienceGuru: an AI improving how LLMs are trained, the core loop of recursive self-improvement. ScienceGuru already runs that loop on the training of our own Guru models. It's our fourth public RSI result this month, after #1 on Autoresearch@Home, the validated lead on @MedARC_AI 's NanoPath v2 and 24.90 s on NanoGPT Speedrun. Not bad for a tiny Singapore startup. Caveats: one run (seed 42), self-reported, not yet on the official leaderboard. Code, logs and source hashes are open, and we'd welcome independent reproduction. Built on @karpathy's nanochat and Giovanni Zinzi's open recipe. Thanks to everyone pushing this benchmark forward. Code: github.com/AutoTrustAI/gpt2-… Try ScienceGuru: scienceguru.ai
2
6
231
What happens when an AI research system picks up where two years of human optimization left off? The benchmark is the NanoGPT Speedrun: train GPT-2 to 3.28 validation loss on FineWeb, the target set by @karpathy 's llm.c replication, which took 45 minutes to get there. The speedrun's code descends from llm.c's PyTorch trainer, itself descended from NanoGPT, hence the name. Over two years, 91 official records brought the time down to 67.56s. ScienceGuru just hit 24.90s on 8×H100 (five seeds, self-reported): 108× faster than where the benchmark started, 2.71× faster than the official record, and 3× faster than Recursive's June record of 75.4s. Earlier this month it also took #1 on Autoresearch@Home at 0.8895 BPB, ahead of @Recursive_SI 's 0.9109, and the validated lead on @MedARC_AI 's NanoPath v2. We built on open PRs by Deven Pietrzak (ANVIL2) and Herman Brunborg (Exact-match), credited in the repo. ScienceGuru's RSI recipe fused them, cut the schedule from 1,194 to 652 steps and re-engineered the host, for the fastest time we know of on this benchmark. Why it matters: the same loop could become a new kind of LLM training optimiser, an RSI system that researches its way to faster, cheaper training runs. Try ScienceGuru: scienceguru.ai
1
2
4
95
ScienceGuru’s Autoresearch@Home #1 and NanoPath v2 maintainer-validated #1 are both public results. We’ll keep sharing the code, evidence, and lessons as the work develops. github.com/AutoTrustAI
RSI claims only count when they're verifiable, and open benchmarks + open code is the right bar. AutoTrust's ScienceGuru has put up two public RSI results (Autoresearch@Home #1, NanoPath v2 #1 validated), recipes on GitHub. Would love to contribute more to the @OpenRSI community.
1
63
ScienceGuru retweeted
1/ Today AutoTrust AI open-sources JEV, a System 1 model distilled from TypeSafe's Jev 1.13: - Nearly indistinguishable from close-sourced Jev (KL 0.021) - Calibrated probabilities, one forward pass - 2.5 ms per decision, batched - Apache-2.0 weights + code huggingface.co/autotrust/JEV
2
2
4
128
Research is full of repeated, bounded decisions: Does this paper meet the review criteria? Which experiment output needs another validation step? Does a result satisfy a predefined quality threshold? Which follow-up action should the team take? These are not the same as asking a model to write an answer. They require a decision, a probability distribution, and a clear understanding of uncertainty. autotrust/JEV is an open-weights System One decision model built for typed yes/no, multi-option, and 0–5 scoring tasks. It returns calibrated probabilities in one forward pass—useful for building explicit, reviewable decision steps around research workflows. It does not replace scientific judgment. But for well-defined operational decisions, it can help researchers spend less time repeatedly sorting, routing, and scoring—and more time on the questions that require deeper thought. huggingface.co/autotrust/JEV…
Introducing the evaluation results for autotrust/JEV: an open-weights model for fast, calibrated System One decisions. JEV takes a typed question over text or structured state and returns a probability distribution in one forward pass: - yes/no decisions - 2–16 option choices - 0–5 ratings No generation. No JSON parsing. No prompt-engineering loop. On a held-out set of 29,955 questions across 53 domains, autotrust/JEV achieved: - 0.021 mean KL divergence from its teacher’s probability distributions - 0.996 AUROC on yes/no decisions - 89.8% top-1 agreement on all multiple-choice decisions - 95.4% agreement where the teacher has a clear preferred option - 0.0007 expected calibration error - 0.10 mean error on 0–5 ratings It also retained meaningful transfer on task families not seen in training: 0.234 OOD KL, 91.8% top-1 agreement, and 0.989 yes/no AUROC. Full evaluation details, methodology, per-domain breakdowns, calibration results, and limitations: huggingface.co/autotrust/JEV…
2
66
ScienceGuru 3.0.32 is now available for Windows and macOS. This release installs as a separate application. It will not overwrite your current ScienceGuru installation. Important: existing projects, standalone chats, session history, and long-term memory are not migrated automatically. Please keep your previous installation and original data directories until migration and verification are complete. Download ScienceGuru 3.0.32: scienceguru.ai/ Migration guide: crm-assets.autotrust.ai/d308…
3
2
3
51
Billing update complete: ScienceGuru Credits are now USD account balance. Your previous Credits were automatically converted at 180 Credits = US$1. Any remaining GPU balance was added at US$1 = US$1. No action is required. What changes: • Model usage is now charged directly in USD at each model’s official API rate, based on actual input and output usage. • You can review every balance change and request charge in Billing Records. What does not change: • For the same usage, your actual cost does not change. • Your Guru Plan price, weekly allowance, and benefits remain unchanged. • Guru allowance is always used first for eligible Guru models. Your USD balance is not charged unless the allowance is exhausted and you choose to continue with account balance. Review your balance, Billing Records, and “Use account balance after allowance is exhausted” setting: scienceguru.ai/profile#/guru Model information: autotrust.ai/models
8
Before removing the previous version, verify: • Required projects open correctly in 3.0.32 • Project files, documents, and experiment outputs are accessible • Required session history has been imported • Long-term memory can retrieve information from prior work • Each standalone chat has been migrated separately • At least one non-destructive functional test has completed successfully Need help? Visit scienceguru.ai/ and use Support.
6
Migration takes three steps: Install ScienceGuru 3.0.32 and sign in with your existing account. Open the same project root folder used in the previous version. Do not select the old ScienceGuru application, installation directory, or an individual file. In the opened project, import session history and rebuild the long-term memory index. Before starting, retain the old installation. Confirm that .scienceguru/ and .ccm/ remain in the original project directory.
18
Local inference matters when research context gets large. For long-running research and agent workflows, model quality is only part of the equation. You also need enough local memory for context, tool use, and multiple active sessions. AutoTrust’s new GLM-5.3 Flash GGUF build uses SLIM-Q to run fully resident on a single 128 GB DGX Spark: 79.1 GiB of weights, with room for 64K context or 4 × 16K sessions. A practical local option for chat, coding, and agent workflows.
Running a capable MoE locally is often a memory problem before it is a compute problem. GLM-5.3-Flash-GGUF-DGX-Spark is AutoTrust’s single-DGX-Spark build of GLM 5.3 Flash.@Zai_org It uses our SLIM-Q compression recipe to reduce the GGUF footprint from 90 GiB to 79.1 GiB — 12% smaller — so the model can stay resident on one 128 GB DGX Spark with room left for context and multi-session work.@NVIDIAAIDev @nvidia What that headroom enables: • 64K context comfortably, or 4 concurrent 16K sessions • Local OpenAI-compatible serving through llama.cpp • Continuous batching, tool calling, and separate reasoning_content output Built for local chat, coding, and agent workloads. huggingface.co/autotrust/GLM…
1
2
76
Research rarely fits into one prompt. A real project can involve a growing literature review, project files, analysis scripts, experimental notes, tool outputs, and a draft that keeps changing as evidence changes. That is why local capacity matters. More memory means more room to keep research context and active work in motion—rather than constantly reducing a project to fragments. ScienceGuru is built for the full research loop: from discovering promising directions to designing rigorous experiments, reasoning through evidence, and developing the work into a written output. Your research workspace should stay yours. scienceguru.ai
15
Turbo 1.2 Preview is now live on ScienceGuru Web. ⚡ Guru Plan subscribers can try it free from September 15, 17:01 to September 19, 17:00 Singapore Time. From September 19, 17:01 to October 1, 17:00 Singapore Time, Turbo 1.2 Preview usage is 50% off. Log in, select Turbo 1.2 Preview from the model picker, and put it to work. scienceguru.ai/
1
2
2
54
Turbo 1.2 is built for real deliverables, not just answers. Work across code, research, documents, and multimodal tasks with faster responses and more stable long-task execution. That direction is already tested in public: ScienceGuru + Guru Turbo 1.0 reached #1 on Autoresearch@Home and #1 on NanoPath v2's maintainer-validated trainable track. 🔬 Your real tasks and feedback will help shape what Turbo 1.2 becomes. github.com/AutoTrustAI/nanop…
22