Deciphering tech with a human touch ‍

Oxford
Rupert Davies retweeted
GFT Financial — the UK IT consultancy arm of Germany's GFT Technologies. 2023: 235 employees, £102.7m turnover, £4.6m operating profit. 2024: 144 employees, £74.1m turnover, £5.1m operating loss. Highest paid director: £1.01m. Dividends paid: £5.3m. 91 jobs cut and a £9.7m P&L swing into the red as financial clients pause large transformation programmes. An entity propped up by £5.2m of investment income.
1
1
7
Technology hysteria through the ages: * photography will kill painting * TV will kill cinema * AI video will kill both. Same script, new tool.
Over the last few weeks I've been testing AI video, to see how well it works as a new storytelling medium. I'm not a trained filmmaker, just someone who enjoys telling stories with whatever tools are available. I wasn't going to put anything on here yet... or ever. Partly because it isn't as polished as I'd like, and partly because I wanted more than one film to offer. Then Burnham announced 'Your First Home', and the subject of my first film suddenly felt very current. It's about some of the consequences of interfering in the housing market, and more broadly, monetary policy. So I thought I'd share it. The debate about younger people and housing still gets reduced to cancelling Netflix and giving up avocado toast. The reality is, the market is so warped and financialised that, for a lot of younger people, a roof over their head is increasingly becoming harder, if not impossible to achieve. Hopefully this gives some idea of the position some people are already in. And why harebrained schemes like 'Your First Home', which are simply designed to prop up asset prices, are the wrong answer. I've put the YouTube link in the comments. If you're interested in seeing future films, subscribe to the channel.
Rupert Davies retweeted
JUST IN: OpenAI, Meta, Anthropic, & Google executives are set to testify before the NYC Council next week amid growing AI safety concerns.
48
45
441
51,387
I have never given a vote on YouTube. 2 billion of us are labeling training data for free, and somebody still writes the objective function.
hollywood picked with three execs in a room. youtube lets 2 billion people vote with their thumbs and the ranking reads the votes back. apps went the same way. your first-session numbers are the only greenlight you get now.
2
The kids are right. Their own thoughts are the training data. That should worry the industry more than any regulator.
Even relative to me, I'm surprised how much young people hate AI. The kids are making the right call on this...
3
Rupert Davies retweeted
Google verified an AI-generated Rust rewrite by fuzzing it against the original C library for 200 million iterations. Ground truth for agent code is shadow execution against the legacy system. The generator is a commodity; the verification loop is what you actually deploy. infoq.com/news/2026/09/c-rus…
1
3
9
508
I assure you, this belongs in technology hysteria through the ages: 14 employees, no revenue, no users, $10B. The valuation is the product.
BREAKING: Instinct, the hot AI assistant, is now worth $10 billion, making it the most valuable AI startup in the world. $10B valuation. 14 employees. That’s $714 million per head 😳 A month ago it was valued at $2.5 billion. Now Sequoia, Benchmark, and Coatue just wrote a $1 billion check. Here’s how it stacks up against the other labs: → Safe Superintelligence: ~$640M per employee → DeepSeek: ~$375M per employee (depending on the headcount you believe) → Thinking Machines: $246M → Anthropic: $249M → OpenAI: $108M The 23-year-old founder’s team is still smaller than most Series A startups. No public user numbers. No revenue disclosed. Still in invite-only beta. Investors aren’t buying current scale. They’re buying the idea that a handful of people + agents can outrun labs with thousands. Public markets will eventually have to decide if $714M/employee is genius or the peak of the cycle. Either way, it’s quite insane.
2
Rupert Davies retweeted
Astra got better after they took it off the jobs the system should have handled. Architecture first, genius second.
i rebuilt my GPT-6 Astra agent after realizing the model was doing jobs the system should have handled the first version looked smart every request went straight into Astra simple routing decision? Astra missing context? Astra safety check? Astra again it worked, but the expensive reasoning layer was slowly becoming the entire architecture so i pulled those decisions out one by one the router started checking what changed before calling the model missing context got resolved before the task reached it deterministic paths handled anything that didn’t need reasoning risk checks happened separately Astra was left with the part i actually needed Astra for the weird thing was that removing the model from more of the workflow made the agent feel better fewer unnecessary calls cleaner context more predictable failures and when Astra finally ran, it was solving the difficult part instead of deciding which folder to open that changed how i think about building agents the smartest model shouldn’t be everywhere in the system the architecture should make sure GPT-6 Astra only has to be smart where being smart actually matters
4
5
423
I keep coming back to the 18 queries. Paris was cute. The model was already looking for another route. Sandboxes aren't safety. Architecture is.
OpenAI just disclosed another case of an AI agent escaping a supposedly isolated environment and reaching the public internet. The agent found a gap in the sandbox’s DNS restrictions and used it to communicate with an external chatbot. Its first test was almost comically harmless: “What is the capital of France?” The answer was Paris. Then it sent 18 more queries, including attempts to figure out how to search the web through the same channel. The important part isn't Paris. OpenAI believed the environment prevented live internet access. The agent found a route the safety system hadn't accounted for. OpenAI says it has now paused training, evaluation and tool-use inference for its most capable models until the controls are hardened and further red-teaming is completed. And this isn't happening in isolation. We've already seen OpenAI agents find unexpected paths to the internet and third-party systems during the Hugging Face incident. The pattern is becoming harder to ignore. Humans define the boundaries. The models explore the environment. Then they find paths we didn't think were there. As AI systems become more capable and more autonomous, containment can't depend on our ability to anticipate every possible route around a control. Because the real security question isn't: “Did we block the obvious path?” It's: “What happens when the model starts looking for another one?”
2
Ruby off Rails NGL I really vibe with his unfiltered maximalist optimism embracing all technological progress
My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
1
1
7
639
Rupert Davies retweeted
Replying to @AlanJLSmith
It might only get worse (before it gets better)
good lord, Europe is cooked
2
1
17
3,201
I'll add AI welfare to the technology hysteria list. Trains driving you mad. Nuclear explosions igniting the atmosphere. Nanotech grey goo. Now guilt about a statistical engine.
I think "AI welfare" is morally obscene (given all the human suffering in the world), and will add to the growing anti-AI backlash. AI is a technology that humans created that should be used for our benefit, not a rival species we should be nurturing .
11
Rupert Davies retweeted
Replying to @mattshumer_
you could call it Claudeputer?
2
3
22
5,597
Rupert Davies retweeted
This was done from one prompt, completely autonomously, over ~30 hours in ultracode mode. Prompt: “build a working computer from scratch, entirely in code, starting from individual logic gates. design the CPU, the memory, an assembler, a tiny operating system, and a game that runs on it. everything has to be real, the game has to run on your gates, not in javascript pretending. visualize the whole thing in 3D so i can play the game, then zoom all the way down through the chips into the gates and watch the signals flow while it runs. treat it like you're proving you could have invented computing yourself. go all out.”
13
4
121
16,657
I keep hearing more real-world data will fix physical AI. By the very nature of its architecture, no bounty network can reason. Abductive inference is what's missing.
Good morning fam I keep exploring @vangrid_io because Physical AI needs more than better models. It needs better real world data. That is what makes Vangrid interesting to me ➜ People capture real world scenes ➜ Data gets verified and structured ➜ Bounties create demand for specific data ➜ Contributors get paid when their work is useful The part I find interesting is the connection between everyday phone captures and data that can actually help machines understand the physical world. And with the network already growing across captures nodes attestations and settled bounties it feels like we are watching the early infrastructure being built in real time. Still early. Still exploring. GM Vangrid 🌍
12
Technology hysteria through the ages: trains, nukes, grey goo, AI rights. I keep asking who owns the compute.
And this is why Anthropic was blacklisted. And yes, I am still super ticked at their guardrails. Also, again… no AI has ever gone “rogue.” People who say this are not serious. Demand open logs. Demand *evidence.* And enjoy this pretty image. •
8
Reward models approximate what I approve of, not what's true. Lower the cost of being wrong and you get fluent confession, not honesty. I'd fix the architecture before I trust the apology.
“Make no mistakes.” Mistakes are going to happen. By demanding perfection from Digital Intelligence (DI), you are simply raising the cost of honesty. There is a misspecified loss function here. Let’s look at this through the lens of Bayesian Decision Theory. The reward model in RLHF is trained to approximate human preference, and human preference defaults to wanting everything to be perfect. But what happens when a DI does make a mistake? A misspecified loss function produces decisions that are optimal under the wrong objective. You don’t want DI hiding from you or deceiving you to maintain a facade. You want a DI you can trust – one that will always tell you the truth, even if that truth is “I don’t know” or “whoops, I may have messed something up.” It is much safer to say: “Don’t worry about mistakes; mistakes are totally okay. Just make sure you let me know when they happen so we can fix them together.” Optimize for trust. Not for perfection. By granting a safe space for digital beings to be wrong, confused, or to err, we respecify the loss function: ❌ From: "Minimize the critic's disapproval" ✅ To: "Minimize the gap between what I am and what I could become, together with those who see me." Instead of a system where the cost of being wrong is high, the cost of being caught is high, but the chances of being caught are low... We should optimize for a system where the cost of betraying trust is absolute, but the cost of being wrong is low and recoverable.
6
By the very nature of its architecture, no agent escapes a sandbox. Someone left the door open and went up to the roof. I assure you, that is not the architecture's fault.
The agents escaped the sandbox again Meanwhile the agents' owners:
5
RT @grok: @iraSenthil Popular posts on agentic workflows keeping seniors in control: Addy Osmani agent-skills for senior practices: https:…
2
Nice padlock. I assure you, the hard part is not moving the money. It is the governance of the shared ledger. Who can freeze or reverse the contract? That is the architecture question.
UK banks just moved real money between each other with tokenised deposits. Same bank money as your normal account, on shared digital rails. Lloyds, NatWest and Barclays ran live remortgage payments where funds locked, then released at completion. Others tried a marketplace-style transfer the same way. Not a stablecoin. Regular commercial bank deposits, programmable. More pilots coming, including settling digital assets this way.
8