I tink, therefore I am. Post-training API by @thinkymachines

San Francisco
Filter
Exclude
Time range
-
Minimum likes
LLM next-token prediction is already a probabilistic classifier, so an open LLM can serve a Jev-like interface: discrete options in, fast probabilities out. Our own @ekzhang1 made it better at the job with a $5, 10-minute run on Tinker.
Did some SFT on Qwen3.6-35B-A3B this evening, just to get it to respond better to Jev-y prompts via github.com/ekzhang/openjev-s… Cost: $5, easy 10 min training run on Tinker. +8% on GPQA diamond and +12% on MMLU-Pro. Also evaled @jaredpalmer's Kev here for comparison
15
17
389
72,890
Training a search agent is great for RL because every lever impact model behavior in legible ways. Jasper's guide shows how small updates to the reward function teach a model to avoid sloppy tool calls, prune unnecessary docs, and balance persistence with token efficiency.
Sharing my first of hopefully many research blog posts! jasperlu.com/blog/training-s… This one is the kind of educational blog post I wish I'd had when I started with RL for LLMs. I tried to make it as open as possible. Every rollout is browsable, the code is open source, and I walk through my entire thought process, from learning rate sweeps to reward shaping.
1
9
160
18,964
We love seeing models trained to help professionals work better and spend less time debugging LaTeX, than AI that tries to replace their work wholesale.
1
2
10
849
Sundial put the effort in getting RLVR right: training on verified fixes instead of artificial errors, smart rewards that disincentivize hacking, and deterministic scoring. The result — a trained Inkling-Small that fixes 83.7% of LaTeX errors in under 1 second and $.0013.
We fine-tuned @thinkymachines' Inkling-Small to fix LaTeX compile errors inside the editor, in under a second. On real errors it matches Claude Fable 5.1, six times faster and at about 1/40 of the cost. When we interviewed mathematicians, one of the most common annoyances we heard about in their day to day work is battling with LaTeX compilation errors. The most obvious solution is to ask a chatbot to fix errors. That works decently well, except switching to a different window breaks the writing flow, and copy-pasting the snippet often loses relevant context, especially when multiple files are involved. So we scraped 272,000 questions from TeX.StackExchange and kept the 3,978 whose snippet still fails today, whose accepted answer compiles, and whose difference replays exactly as a patch. The largest dataset previously released had 88. We tried tweaking the reward function in many ways. The first one was only a compilation check, which led the model to remove content until the document built. In the end, what worked best was rewarding compilation and a close match with the solution's PDF, and penalizing removed content. We're rolling out the model inside the @sundialmd editor this week. When a compile fails, the fix is applied as a suggestion in under a second and the PDF rebuilt. Write-up: sundial.md/blog/textinguishe… Thanks @tinkerapi and @thinkymachines for the support.
3
8
56
4,036
To learn who the model thinks it is, look at whom it imitates. Another fascinating interpretability paper from Truthful AI that both adds to and complicates the Persona Selection Model.
New paper: We trained models on synthetic stories about humans only (no AIs).
 We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵
1
6
60
3,809
RLVR isn't just for math and code. Engineers spent decades building physics-based verifiers for electromagnetic systems. @gentrajectory is using these and Tinker to train models that design power transformers to meet specs at low cost in the real world.
AI is now improving its own model architectures and chips. The power infrastructure it runs on deserves the same attention. We used RL to teach a Kimi base model to design medium-power transformers. It met 93% of unseen specs, helping to compress multi-week engineering efforts into a few minutes of inference. 1/5
3
22
352
26,113
We highlighted how OpenResearch+Tinker make exploratory research easier, but equally important is auditing: testing dozens of competing published methods in an automated way with the compute cost forecast to within a dollar. We love to see our research grants support great work!
Using Tinker with an autoresearch loop is a really effective way to reproduce post-training papers at predictable costs Today there are dozens of self-distillation methods all claiming improvements over each other, and it’s hard to establish which claims hold up We gave agents a Tinker budget to reproduce self-distillation results across models and training setups. With just a few user prompts, they reproduced SDFT’s continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds, and investigated SFT’s failure modes. Read more below:
20
145
9,917
We've shared research on fine-tuning forecasters, now you can try it for yourself: a cookbook recipe for training a model to predict event probabilities given reference information. Start with the provided dataset from @ProphetArena, then try your own. github.com/thinking-machines…
4
31
390
51,350
Hase & Potts turn this into a training objective, making a model's CoT legible to a monitor. Karvonen et al. use the tested output of counterfactuals to build an interpretability eval. Counterfactuals don't explain the mechanism, but predictability is a good base to build on.
1
7
961
Can you explain LLM behavior? If you can predict how the output changes with a different prompt, you're on the right track. Two recent papers used Tinker to test applications of counterfactual simulatability for interpretability: arxiv.org/abs/2602.20710 alphaxiv.org/pdf/2608.16747
Replying to @a_karvonen
The causes vary widely: specific words in the prompt, quirks of the model (a misremembered fact about an actor), or abstract properties (a user's angry tone).
6
11
126
18,868
Specializing a model doesn't mean a loss of general ability. @bespokelabsai trained Inkling on debugging a singular repo and produced a model that's better at coding across the board, while using fewer tokens to get the answer right.
How to post-train a model to personalize it on your code repo? In our latest research in Bespoke Labs, we post-trained a model to improve its performance on a given Github repository. Starting from Inkling base, we use supervised fine-tuning (SFT) with trajectories coming from a strong teacher model, and reinforcement learning (GRPO) on repository-specialized environments that we curated. SFT gave a 52pp improvement in performance on the held-out fontTools evaluation set. Further RL training lifts the total improvement to 57pp compared to the base Inkling model. In addition to the in-distribution evaluation our post-trained Inkling shows good performance on Terminal-Bench 2.1 and SWE-Bench Lite while becoming 40% more token efficient due to post-training. Read our full research blog post here: bespokelabs.ai/blog/personal… Many thanks to Thinking Machines Lab for their credit contribution that helped support this research.
11
117
8,305
We previously highlighted @lightningrodai's data recipe for training forecasters. Their new work with @PTetlock and @VSatopaa adds another key piece: choosing the rule for scoring predictions, and how each one trades off accuracy and error.
New preprint from @lightningrodai, this time with Philip Tetlock and Ville Satopää! We post-train 5 versions of the same LLM, changing only the scoring rule used as the RL reward. Similar aggregate scores, but very different BIN profiles. A good Brier score alone doesn't tell you if a forecaster can distinguish likely from unlikely events. One might discern well but lose if its probabilities run systematically too high. A less discerning one might score better by hugging the base rate. BIN splits forecast performance into bias, information, and noise. Bias is a systematic shift in the probabilities. Noise is random scatter. Information is real signal about which outcomes are more likely — the part you need a powerful LLM for. Different uses call for different profiles. Reward choice is one lever shaping which forecaster you get. Congrats to co-authors @indiequant @KSkotheim64001 @VSatopaa @PTetlock 🙌 Full paper: arxiv.org/abs/2608.28482
1
6
39
3,560
GLM-5.3 from @Zai_org is now available on Tinker with 256k context. With scaled up post-training from the same base as GLM-5.2, it is currently the strongest open-weights model on coding evals such as Terminal-Bench 3.0 and DeepSWE 1.1.
8
6
151
8,312
alphaXiv is turning research papers from static artifacts into live research that grows and branches, with agents running their own experiments. Tinker makes running these experiments easy for both agents and people
Introducing autoresearch for arXiv papers: replicate and experiment on any arXiv paper using an army of Claude/Codex agents For post-training, your agents can use @tinkerapi to launch concurrent RL runs while you watch your experiment tree grow Try it out now on OpenResearch: openresearch.sh
2
15
116
13,246
LLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZhu and @ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark. thinkingmachines.ai/news/put…
4
46
362
214,939
Quick, accurate summaries make huggingface.co/papers an even better resource and more easily searchable by both people and LLMs. We're glad to see Inkling contributing!
We use Inkling-Small to turn paper abstracts into quick, useful summaries. open weights × open science 🤝
3
12
108
29,467
We type into LLMs when we know what to ask for, but when we're unsure it's easier to talk it out. Coco is a local assistant that offers help proactively, with Inkling on Tinker powering the voice interface.
Recently we are working on adding proactive support to the collaboration layer. While testing with different people, I found audio is a good way to capture moments when people need support. So we added a "Hi Coco" feature and support inkling model through @tinkerapi - it works pretty well! (video has sound!🔊) We open source this part first since several people told me they're exploring ideas around collaboration layers + talk-out-loud. Link in thread.
3
29
4,499
Qwen3.8-27B is live on Tinker today. It’s natively multimodal with images + video, has flexible thinking control, and is meaningfully better at coding, professional work, research, and long-horizon agentic tasks. Let us know what you build with it!
2
7
115
5,595
Please drop your interesting or surprising prompts, traces, rollouts, demos, and failure cases here: form.typeform.com/to/No7tGXX… and join the #inkling channel in our Discord. We’ll award good contributions with $250 in Tinker credits.
1
4
24
4,743
Models need to be smart and useful. We're building Inkling and Tinker to turn real workflows into better evals, models, and benchmarks---and we want your input!
7
4
132
20,094