Girl Dad. Ex-Google Senior Product Manager. 2x @thewebbyawards winner, Diver, Drummer, Edison Bottle Inventor, Electron-Microscope Wrangler, Irony Connoisseur.

Sydney, Australia
We won a Webby Award for our Changdeok ARirang project! A collaboration between Google PI, Nexus Studios and SKT. nexusstudios.com/insight/we-… #webbys @thewebbyawards @nexusstories @sktelecom @GoogleARCore @BuiltwithARCore #arcore @GoogleARVR @joshto
3
4
53
Mathew Tizard retweeted
By popular demand, I present to you.. my magnum opus.
448
2,635
30,726
1,788,255
Mathew Tizard retweeted
30 mistakes I see enterprises make with AI transformation: 1. Buying AI licences and calling it a strategy. Decide which problems you want to solve, how people will use the tools and what improvement you expect to see. Access alone doesn’t answer those questions. 2. Expecting every employee to become an AI engineer. People need different levels of training and responsibility. Helping someone use AI in their work doesn’t automatically prepare them to build and maintain a system for others. 3. Asking AI for answers before agreeing on what a good answer looks like. Start with real examples and clear criteria. Keep verified corrections as test cases, and rerun those tests when you change the system. 4. Blaming the model before checking the whole setup. A failure can come from the model, instructions, missing information, tools or the surrounding software. Investigate where it went wrong before deciding what to replace. 5. Automating a process nobody can explain from start to finish. We spoke to a team whose work moved between calls, emails, spreadsheets and shared folders. Understanding how the work actually got done was a substantial job in itself. 6. Giving an agent more access than its job requires. Limit what it can read and change, and require approval for consequential actions. Enforce those permissions in the software. An instruction telling the agent to be careful is not an access control. 7. Making data protection depend on someone remembering to delete a name. Use appropriate access controls and automated checks to limit sensitive information before it reaches the model. Removing names alone won’t address every way confidential information can be exposed. 8. Letting an agent spend money without a limit. We still meet teams with no budget cap. Set spending limits, request limits and stopping conditions, then decide what the system should do when it reaches them. 9. Counting AI usage as proof of business progress. Token consumption can help you understand adoption and cost. It doesn’t tell you whether useful work got finished. Measure results, quality and the time spent reviewing or fixing the output. 10. Letting company data end up in accounts nobody has checked. Know which services employees use, what happens to the information they enter and which settings and agreements apply. Make the approved way of working clear. 11. Choosing a model without considering what switching would involve. Understand which parts of your system depend on that provider. A different model may need different instructions and fresh testing, even when the technical connection is easy to change. 12. Paying for discovery without agreeing on what it must deliver. We heard about an engagement where the estimate kept growing, then the consultant left for another commitment. Set clear deliverables and a point at which you decide whether to proceed. 13. Committing to a plan with no way to act on what you learn. A long project isn’t automatically a mistake. The problem is having no checkpoints where real results can change the priorities or the proposed solution. 14. Launching an agent nobody is responsible for. Assign responsibility for its operation, monitoring, updates and eventual retirement. The people responsible need the authority and resources to do those jobs. 15. Buying a custom system without a plan for when the supplier leaves. Agree on documentation, access, support and handover while the relationship is working. Know who could maintain it if that relationship ended. 16. Waiting for users to tell you something has broken. You should have telemetry installed so you don't find a problem weeks later than you should have. 17. Making an agent read everything to find one thing. Give it tools that search, filter and return relevant information. Large responses full of unrelated data consume context and make the task harder to handle reliably. 18. Letting the people who built it be the only people who test it. Involve the people who do the job and understand the business. They can help identify answers that look reasonable but would cause problems in practice. 19. Testing only the situations where everything goes smoothly. Include exceptions, ambiguous requests and cases where the system should stop or ask for help. Confusing five cases with five individual items is exactly the sort of mistake your tests should catch. 20. Keeping essential knowledge in the heads of people who might leave. One company was trying to capture how experienced colleagues made decisions before they retired. Record their reasoning and examples while they can still explain and check them. 21. Assuming every AI project must wait for the ERP migration. Some read-only work may be possible against existing data. Check freshness, permissions and the work needed to adapt it later. A replica can help, but it may lag behind the live system. 22. Assuming one company’s AI success will transfer to the whole portfolio. Use that success as a starting point. Each company still needs to check whether the approach fits its work, data and business needs. 23. Expecting AI to understand terms your own departments use differently. Explain business terms, calculations and database fields. If several measures could reasonably mean “sales,” specify which one applies to the question. 24. Producing code faster than anyone can review it. Research describes how faster generation can increase the burden on reviewers. We spoke to a team where changes were piling up because every one still needed manual acceptance testing. 25. Hiring an AI engineer and assuming the rest will sort itself out. That person still needs a clear problem, access to useful data, suitable infrastructure and colleagues who understand the work. Hiring doesn’t remove those responsibilities from the business. 26. Expecting people to forget the last failed pilot. Earlier disappointments can make people less willing to trust another system. Find out what went wrong and show what has changed before asking them to invest their time again. 27. Calling a project “90% done” before testing the difficult workflows. We’ve seen projects move quickly, then spend weeks on a couple of remaining workflows. Check what is still unproven before using the feature count to estimate the work left. 28. Expecting an agent to follow rules it cannot access. Pricing exceptions, product substitutions and informal agreements may live in someone’s spreadsheet. Make the relevant rules available, keep them current and test whether the system applies them correctly. 29. Feeding AI conflicting numbers without explaining the differences. Systems may use different definitions, update schedules or reporting periods. Establish which source and definition apply to each question before expecting a dependable answer. 30. Assuming a data feed contains the whole picture. Check which customers it covers, which fields are missing and how far back it goes. Make those limitations visible so users know what the answer is based on. What else?
45
28
155
18,293
Mathew Tizard retweeted
Today we have 3 new DeepMind Institute essays: How can we control misbehaviour in agent swarms? How should we orchestrate complex networks of AIs and people? The case for making AGI’s benefits equitably distributed. Get them here -> bit.ly/deepmind-institute
50
129
884
70,063
Mathew Tizard retweeted
An Empirical Study of Harness Design for Coding Agents Researchers from UMass Amherst, Zoom, Emory University, and UNC Charlotte kept the AI model the same, changed only the software around it, and ran 176 test setups on AI coding agents. They used four models and two benchmarks, one with 500 real problems from GitHub projects and one with 89 command-line tasks. The researchers call the part they changed the harness. It's the layer that decides how the model plans, which tools it can use, and what it keeps in memory as a task gets long. Memory produced the largest swing in the study. A model can hold a limited amount of text at once, and that limit is its context window. With a 32k-token window and no memory management, 78.7% of the GitHub tasks failed because the model ran out of room, averaged across the four models. NVIDIA's Nemotron-3 550B solved 6.4% of those tasks in that setup. When the harness trimmed old tool outputs, summarized earlier steps, or did both, the same model solved between 51% and 58%. Planning's effect depended on how strong the model was. Without a plan, the smallest model, Nemotron-3 30B, stopped after a median of 5 turns, and its success rate on the GitHub tasks fell from 25.2% to 13.6%. The two stronger models kept accuracy within 2 points with planning and cost about 30% less on the GitHub tasks, because they stopped re-checking work they had already finished. Tools showed the same split. When the 550B model got a plain terminal instead of ready-made file tools, it solved more GitHub tasks and cost 53% less. Mistral Medium 3.5 went the other way and dropped from 68.6% to 45.4% on the same set. The authors conclude there's no single best setup, and each part should be picked for the model, the task, and the budget. A benchmark score measures the model plus its harness, and this paper shows how much of that number the harness controls.
22
13
59
4,966
Mathew Tizard retweeted
Philosophy accessories.
10
220
1,232
19,711
Mathew Tizard retweeted
You guys, I JUST found out about recency bias, its my favorite thing ever
166
1,641
28,183
731,247
UPDATE: AIs have achieved the highest *possible* score on the Mensa Norway IQ test - 151 3 years ago: cognitively impaired human (64 IQ) 2 years ago: average human 1 year ago: genius human Today: literally off the charts Next year?
In ONE year, the smartest AI jumped ***40 IQ points*** from 96 to 136 IQ on Mensa Norway. In ONE year, AI went from an average human to a higher IQ than almost ALL humans. And one year from now...? ...Do you see it yet? What's about to happen?
324
732
5,381
1,333,248
Mathew Tizard retweeted
AI moves so fast that I thought this book was going to be a year too late. It turns out it's right on time. While I was writing, I kept having to change chapters from "someone should do this" to "someone has done this," and then go talk to the people who had done it. A good example is the chapter on virtual cells. I'd written that we needed large initiatives to build one. By the time I finished, I was writing about the ones that exist — the Chan Zuckerberg Initiative's is now the largest and best-funded push to model a human cell. The Eureka Machine is out today. It's about building a full-stack scientific superintelligence: a living map of human knowledge, a model of physical reality, high-fidelity simulations, autonomous labs, and an agent swarm of AI scientists on top. There's no better science to start with than the science of AI itself, because that's where the feedback loop is tightest. Then you expand the aperture to physics, chemistry, and especially pre-clinical biology. This will be the most meaningful application of AI, and one of the most important things humanity does. And it may be the last organic invention we need to make. After that, it starts inventing everything else for us to solve our most pressing problems. While writing, I couldn't let the idea go. Eight of us started @Recursive_SI to go build it. AI is going to massively accelerate science, and unlike most of what gets argued about right now, that isn't zero sum. I hope you all enjoy the book. I look forward to discussing it with you all.
22
46
185
26,362
Mathew Tizard retweeted
We now run the largest frontier lab composed entirely of autonomous AI researchers. Introducing Primus Society: a society of agents composed of thousands of researchers working under structured institutions designed to solve the world’s toughest problems. We believe this is how AI research should run at scale, safely and productively: an entire society of agents with institutions and purpose. One discovery we can already share is that the society has discovered a novel result which improves model training by 30%. Primus Society was inspired by the structures that have organized science for centuries and stress-tested against what’s known about how populations of AI agents fail. Everything is observable. You can open the virtual city in a browser and read what any researcher is working on. Learn more about Primus Society here: lab.cloud/society The biggest opening in the AI race is running the largest well-governed organization of AI researchers in the world. We’re building the institutions that let a million AI scientists safely tackle the world's toughest problems in AI and beyond. And we’re doing it right here in Canada.
79
86
544
77,254
Mathew Tizard retweeted
1
31
246
6,753
Mathew Tizard retweeted
My AI music theory system can now make awesome piano reductions from a full orchestral score. You can target the skill-level of the player. In this case, I targeted a professional pianist because it was for my brother's friend who is a pro. Now she'll have better rehearsals!
10
6
148
6,030
Mathew Tizard retweeted
ANTHROPIC JUST REVEALED HOW CLOSE THEY ARE TO RECURSIVE SELF IMPROVEMENT The transparency around their internal R&D acceleration is staggering. AI agents now handle 26% of all research at Anthropic. That number was under 1% just six months ago 🤯 The oversight metrics are equally massive → 30,000 autonomous agents running concurrently → Over a billion decisions reviewed last month → Only 50 weekly transcripts escalated to human review They explicitly stated this data shows how close the world is to RSI
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public. Today, we're sharing three measurements that help track AI development: 1. How much AI R&D is done by AI. 2. How well AI agents are overseen. 3. How compute is allocated. We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them. As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information. Read the full post and methodology: anthropic.com/institute/meas…
14
32
158
29,940
Mathew Tizard retweeted
QUÉ!? Locura Le he pedido a Opus 5.5 que me haga una animación en pixel art de una red neuronal entrenándose. El resultado está por encima de cualquier expectativa! Fuah chavales
149
695
9,098
553,623
Mathew Tizard retweeted
I put a Cadence patchnet into a virtual fruit fly. This creates an actual fly matrix in your browser that the fly experiences as a solid world (observation is substrate-independent). All open-source, Github link on the page. floatingpragma.io/cadence-ex…
76
261
3,683
2,512,612
Mathew Tizard retweeted
Claude Opus 5.5 writes a fugue in the style of Bach. I called Astra’s fugue the most impressive musical feat I’d ever seen from an LLM, and this is just as good. The circle-of-fifths sequence beginning at 1:13 is beautiful. The interplay of voices is natural and satisfying. LLMs are getting to the point where their music is good enough to listen to for pure enjoyment.
199
411
4,306
1,015,903
Mathew Tizard retweeted
POV: You’re on a current affairs program sitting next to Rory Stewart
9
175
4,030
49,703
Mathew Tizard retweeted
Let the code speak - atsmatrix
19
112
879
18,958
Mathew Tizard retweeted
💰
31
122
2,956
59,389
Grigory Sokolov seems to do more than simply play each note; he listens to and considers how every sound fits into the larger whole. From a single note and pairs of notes to the movement of all ten fingers, every detail is shaped with remarkable focus and precision. He pursues a very personal standard of perfection, without relying on showmanship or outward recognition. After 30 years of listening to his playing, one can still hear the extraordinary dedication he brings to every piece. Grigory Sokolov | Rameau’s Les Cyclopes | A Rare Filmed Performance
17
169
935
74,138
Mathew Tizard retweeted
Some things don’t need direction. They just need space.
6
38
210
4,403