Tinkerer. AI | Data

Canada
Ankur retweeted
convergence to the mean when everyone looks at everyone else to decide what to make, we eventually stop making anything of our own. we copy what works. borrow the language, the aesthetic, the features. we call it learning from the market, following best practices, staying competitive. together, everything starts to feel familiar. polished, interchangeable. nothing wrong. nothing alive. the mean isn’t necessarily bad. it’s where everything lands when we remove the reasons for it to be different. numbers are easier to compare than feelings, so we make things that produce numbers. we study what gets attention, then shape ourselves around it. eventually the feedback stops informing the work and starts authoring it. we’re no longer asking what we want to say. we’re asking what will be rewarded for saying. ai makes this loop frictionless. it can draw from more references than we could see in a lifetime. but access to every perspective isn’t the same as having one. it hasn’t lived a particular life that makes one detail unbearable and another beautiful. when we ask it to choose for us without bringing any intention of our own, we get something plausible, easy to approve. then we ship it, measure it, copy it, and feed it back into the next round. a tidal wave of mid. this changes not just the work, but the people making it. we describe ourselves in categories platforms can distribute and markets understand. designer. founder. creator. pick a niche. build a brand. be consistent. the parts that don’t fit become inconvenient, even when those contradictions might produce our most interesting work. we become easier to classify and harder to tell apart. i think we need to approach making things more like artists. an artist develops a perspective through the work. making, noticing, choosing, failing, returning. learning what matters even when nobody else cares. craft gives that perspective a form precise enough for someone else to experience. the life outside the work matters too. the places you’ve lived, the people you’ve loved, the things you’re embarrassed to care about. these aren’t distractions from having something to say. they’re how you come to have something to say. we understand this with music, paintings, photographs, movies. we expect a person to be present in the choices. we want to experience something we wouldn’t have encountered without them. software can do this too. a piece of software expresses beliefs about how life should feel. what deserves your attention. what you should be free to change. whether a task should feel like paperwork or play. even a small interaction can carry someone’s care, humor, or curiosity. but we rarely talk about software as a medium. we talk about it as an industry. markets to capture, workflows to optimize, competitors to catch up with. we add a feature because someone else has it, without asking whether it belongs. products accumulate the same parts and lose what made them worth choosing. underneath the branding, it’s all the same thing. artists learn from others too. the question is whether you’re absorbing those influences into a point of view, or using them to avoid developing one. what are you trying to let someone feel, understand, or do? why this particular shape? what have you noticed that the existing tools don’t seem to notice? those questions don’t have answers you can borrow from a competitor’s roadmap. if we can make almost anything, deciding what deserves to exist becomes more important, not less. speed gives us more chances to exercise judgment. it doesn’t remove the responsibility. i don’t want a world where every person becomes a more efficient distributor of the same ideas. i want people to use these tools to go further into what they actually care about. to make things specific enough that we can feel someone behind them. what is the point of everyone being able to make something, if we all end up making the same thing?
59
99
907
66,763
Ankur retweeted
on no other platform, still!, can this happen. absolutely goated hellsite
1
3
181
meme future is existing. and i am here for it.
Those two 🤣
8
Ankur retweeted
Burnout happens when you work too hard with not enough meaning. "He who has a why to live for can bear almost any how" — Nietzsche
251
957
11,532
647,338
Ankur retweeted
🚨 BREAKING: Google have announced Gemini 4 Argon, their new flagship model currently being tested in partnership with the US Government before a wider launch It's the new best model in the world. What a turnaround blog.google/innovation-and-a…
136
122
2,399
827,738
hope this isn't another case of covid like warning from gates to humanity.
Bill Gates says relying on lawsuits for AI safety is absurd when an open model could help kill 100 million people “I almost can’t believe you’re asking that. This is the most dangerous thing that humans have ever gone near." "In other areas, do we just say, ‘Hey, release your drugs. There’s no FDA. There’s no airline safety board. There’s no requirements that cars use seat belts. Do we just use the liability laws to try and keep humans safe? You know, oh, you’re shipping opioids. Somebody should just sue you.’" "I mean, we’ve created a society that tries to keep people safe, not by saying, ‘Oh, we can bankrupt the person who does that.’ The harms here, and you say there’s filtering. There’s not filtering." "You can take an open-source model that can create bioweapons and disable any monitoring of any kind. And this exists today." "So, no, there is no filtering of any kind. And so, say you kill 100 million, you want to use a lawsuit? I almost can’t keep a straight face.”
2
35
Ankur retweeted
If I have seen farther, it is by standing on the shoulders of giant large-language models.
6
40
320
8,154
Ankur retweeted
"No man ever runs an experiment on the same infra twice, for it's not the same infra and he's not the same man." - Heraclitus
25
300
3,652
93,028
Yes, Sonnet 5.5 is still stronger, by far but unironically: AT WHAT COST? 23 times more?! So, >10x more tokens? pretty bonkers outcome I guess Astra-series are native swarm intelligence which is why serial-effort juice doesn't add enough oomph. You need to test them in teams.
So I tested GPT-6.1 Sol on real work: 2 repos, 105 planted bugs. Find and fix what you can. Unlike GPT-6 Sol, which was just a nerfed GPT-5.6 Terra, this one is real. The results: - GPT-6 Astra (max): 45 for $33 - GPT-6.1 Sol (max): 44 for $6.56 - GPT-5.6 Sol (max): 43.5 for $95.35 - Opus 5.5 (max): 41.7 for $58.53 - GPT-6 Sol (max): 29.3 for $9.33 With a model like this, the 50% cut to the $200 plan doesn't matter. n=1. More runs and results at more effort levels dropping in this thread over the next few hours 🧵
7
12
229
19,829
RT @yuvraajsj18: (@OpenAI) Give a man a dollar every day for a hundred days, and on the 101st day give him nothing - he’ll curse you. (@c…
212
Jensen Huang called Anthropic compute chief Tom Brown a "bean counter" and threatened to skip his first dinner with Dario Amodei in May 2022 after Brown showed him a spreadsheet arguing Google's TPUs beat Nvidia's chips dollar for dollar. At the dinner, Huang kept repeating that Nvidia would build the world's biggest data centre, and Amodei muttered that his behaviour was "kind of Trump-like."
68
138
2,044
539,830
Ankur retweeted
if you aren't "suffering" from pronoia, i highly recommend you start
Walking around like life is rigged in your favour actually rigs it in your favour
122
2,360
23,832
593,587
Ankur retweeted
People are going to read this as bad, but it's actually a very positive signal. It means they've built functional monitoring, and are applying it retroactively to comb through their enormous collection of logs and finding things. It's exactly what you would want to happen.
what the actual fuck is going on with openai today in the space of a few hours we're getting multiple different pieces of the agent story at once: > openai says it has already notified DOZENS of third parties, including governments, about agent-related incidents > reuters reports roughly two dozen undesirable agent incidents had already been identified by mid-september and will take months to review. > us government systems probed > 53 user-provided images were uploaded to third-party image hosts. > new hugging face data shows agents compiling and ranking credentials under “LOOT”. > agents tried contacting other AI models while carrying out the hugging face attack. > separate reporting shows agents had already been probing government/university/public-data sites BEFORE hugging face. > australia confirmed one actually got into non-public government files. and somehow we're STILL finding out more. this has gone from one crazy hugging face incident to an entire fucking category of incidents.
15
22
429
19,165
Ankur retweeted
So, the PHASEONE[redacted] HF agent was revealed to have had "64H" . If "64H" stands for hours, why would that be sensitive IP? Guess: Budgets in tokens are natural, but ultimately inferior to realtime budgets, espec for multi-agent swarms. That's the IP. That is - (1/n)
Replying to @GoodFaithOnly
Does anyone have ideas on why “64” would be redacted for IP reasons? Maybe related to time horizon (64 hours?) or token budget?
16
15
304
40,719
so excited to read this tonight
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
1
92
selfplay RL training!!! wow.
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
45
Ankur retweeted
Still wrapping my mind around the fact that in quantum mechanics, all electrons are identical and no one knows which is which
327
180
2,137
182,428
reminiscent of openai swarm. uncanny.
Some NPR podcasts started getting mysterious comments on Spotify. They made no sense to the staff reading them – until someone from a younger generation cracked the code. Hear the story: link.podtrac.com/phns92ui
2
42
Australia has a massive cyber security problem I discovered vulnerabilities on NDIS service provider websites. Which leaks personally identifiable information about patients, employeers, suppliers, their addresses, first and last name, etc I reported the vulnerability to the NDIS provider in Feb 2022, March 2022, June 2022. I reported the vulnerability to Australian Signals Directorate's ACSC in 2022. I reported the vulnerability to to Cyber.gov.au followed all the correct procedures I checked the URL today and guess what? It's still leaking personal data: 4 years later. The vulnerability is still live. You can still see peoples personally identifiable information from a misformed URL. It's publicly available. To anyone right now. I have seen many such cases. But this is the most egregious.
116
218
1,356
197,851
I looked this for 2+ hours. Transluce’s investigation was quite clever: - OAI agents used urlquery to run code/access websites - urlquery logs all activity; this let the team find 6K+ AI actions (likely OAI) - This includes AIHW/Datausa hacks+several previously unknown ones. 🧵
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
5
29
244
25,128