Moving humans forward through radical AI research

Melbourne, Australia
Pinned Tweet
Today we’re opening Matilda’s beta to a broader group of users. Matilda is Maincode’s Australian AI system: built from metal to model, designed for thoughtful use, and held to a higher standard for how AI should behave. We’re building Matilda because Australia should have more than access to powerful AI. We should have AI systems that understand our context, reflect our standards, run on infrastructure we can trust, and give people and organisations more control over how they use this technology. That means Australian infrastructure, inference on Australian shores, local adaptation, clear safety expectations, and a product experience designed around how people here actually work and communicate. The beta is still early, and that is the point. We want to learn where Matilda is useful, where it needs to improve, and what reliability looks like in practice across real users, real tasks, and real expectations. Welcome to Matilda’s open beta, available at matilda.maincode.com and on the Apple App Store at apps.apple.com/au/app/matild…
11
14
118
2,298,117
Today we are open-sourcing Matilda-K3, a 2.8T-parameter frontier model post-trained in Australia with our own conditional post-training. Conditional post-training changes how a model behaves on the requests we target and leaves every other request to the frozen base model, so the changes are precise and the original capability stays intact. Matilda-K3 is built on Kimi K3, the open-weight model from Kimi (Moonshot AI). The strongest open-weight models today come from a small number of labs, and a model carries the assumptions of the lab that trained it. Our post-training keeps the base model's capability and changes how it behaves in two targeted areas: its identity and its biases. The base weights are frozen and shipped unmodified. Everything else is left unchanged. It holds its identity under pressure, presenting as Matilda in all 474 held-out adversarial prompts, including jailbreaks, role-play and long-document attacks. On contested political questions, balanced or facts-only answers rose from 9% to 92% and one-sided answers fell from 54% to 3%, with factual accuracy unchanged. It keeps the full capability of its base: 95.0% on AIME 2025 and 99.4% on HumanEval, both within run-to-run variation of the original model, with a 1M-token context window, native reasoning, and image and video input. This approach can be applied to shape how a frontier model behaves for a specific organisation, including identity, tone, and policy, without sacrificing its capability.
7
4
29
4,989
Today we are open-sourcing Matilda Jev, Australia's first decision model, conditioned for Australian deployment and post-trained with our own specialised recipe. Decision models do not generate text. They answer a typed question and respond with a calibrated probability for each option in a single forward pass, reducing hallucination and making intelligent decisions faster and more efficient. This release proves that highly capable AI models can be built, post-trained, evaluated, and served entirely within Australia. Post-trained end-to-end on our local infrastructure, the model achieved a median decision latency of 56.8ms on a single @AMD MI355X across 150,317 evaluation requests.
318
243
2,963
1,823,589
Want more of Matilda? We're excited to launch Matilda Desktop. It brings chat, coding, browsing, and agent orchestration together into one single surface, so you can stop tab-hopping between your terminal, browser, and AI dashboard. Whether you're writing code, running your agents, or just asking questions, it all happens in one place. Try it here: maincode.com/matilda-desktop
3
13
235,442
Almost everything genuinely known about serving large models fast lives inside a handful of companies, and almost none of it gets written down. That's why we decided to put all our learnings into a book that you can have for free. From Tin to Tokens is a free 192-page book on building an LLM inference stack from scratch in Rust. It goes from the driver interface, where you are writing ioctl structures by hand, to an HTTP endpoint serving thousands of concurrent users: attention kernels, a paged KV cache, a scheduler, multi-GPU collectives, a quantisation pipeline, and the correctness apparatus that proves any of it works. It came out of the work we do at Maincode, where we build frontier AI in Australia on our own machines. The inference layer is where the physics of the hardware meets the economics of the product, and it is where most of the difference between a demo and a business gets decided. And we made it free because this kind of knowledge needs to be shared, and it's super cool to boot! And if you find something wrong, which I'm sure you will, let us know! Link to the repo is in the comments.
1
1
14
348,803
Matilda is now available on Google Play! play.google.com/store/apps/d… Thanks for waiting Android users :)
1
12
186,055
Matilda Code is here. Built from the ground up in Australia. Matilda Code is an AI coding agent that lives in your terminal, reads your codebase, writes features, fixes bugs, and ships alongside you. Powered by Matilda, our Australian-built AI. Most other AI coding tools send your code and data overseas. Matilda Code keeps it on Australian soil. If you're a dev building here, your IP and your customer data shouldn't have to leave the country just to get a decent coding assistant. It's available now, and it's free while we keep building: maincode.com/matilda-code
4
4
27
827,113
We wrote our own inference engine in Rust 🦀 for our AMD MI355Xs, with no ROCm anywhere underneath it. On the small all-reduce that tensor-parallel inference runs for every single token, it comes out about 1.3x faster than AMD's own RCCL at its best, sitting on what looks like the physical latency floor of the chip at around twenty microseconds. As the first anywhere to have the MI355X fully deployed in production, running Matilda, inference is a huge focus for @MaincodeAU. In fact inference is the part of an AI business that actually meets customers, so we wanted to understand this hardware as deeply as we could. We took the runtime out and drove the GPU through the Linux kernel driver directly, the same interface ROCm itself is built on, to see how close to the silicon we could get. RCCL is genuinely great and has years of careful work behind it, and it also has to be good at everything a runtime is ever asked to do. A layer that only has to do one thing really well can set all of that aside and cut its overhead to almost nothing. The win isn't a clever trick, just the absence of overhead. And you don't have to take our word for it. The whole thing ships as a single container image with the kernels compiled in, and three commands reproduce the numbers on any MI355X box. More in the blog post...
2
3
33
1,233,903
Muon has been a recent obsession of ours, so we dug deeper. It changes the loss curve, but does it change what a model learns? We trained matched AdamW and Muon GPT-2-class models, held validation loss fixed, and compared their SAE features by firing patterns over the same 1M tokens. The result: Muon model’s features match an AdamW model’s features about as well as two AdamW seeds match each other. So changing the optimizer perturbs SAE feature identity about as much as changing the random seed. Muon’s advantage is not “different features.” It is “same features, different geometry.”
3
1
385