AIUX Research, Design, Strategy, Evals. SF-Aus. IntentCo founder. X @google & @googledeepmind built: @geminiapp Imagen3+,Canvas, Assistant, @googlemaps +

San Francisco, CA
Proud to say I helped build this via foundational R&D for Imagen and Imagen 3 with editing - this is next level !
We've shipped some big updates to the Gemini app over the past few weeks. We launched new creative and collaboration tools like Nano Banana, simplified app creation with Canvas, and more. Here's a look at everything new 🧵
6
35
13,064
KP 🇺🇸🇦🇺 retweeted
holy shit ChatGPT is actually developing humanlike academic skills
thanks chatgpt
86
548
11,522
523,576
KP 🇺🇸🇦🇺 retweeted
Replying to @anah_sahh
These posts have shown me that most people are absolutely clueless as to what is going on with ai and robotics. 99% of the jobs listed here will be 99% ai in two years. Accounting is out in months. You will purchase an affordable robot slave within 12 months from now - and that's me hedging.
17
1
25
4,520
This is a Google Workshop Gemini killer, but, Google does own a large % of Anthropic, so still a horse in their favor
Claude now works inside Google Docs, Sheets, and Slides, and those files also open inside Claude. In Google Workspace, Claude sits in a sidebar next to your file, reads what you have open, and edits it in place. You can approve each edit before it lands.
1
153
KP 🇺🇸🇦🇺 retweeted
I'm not sending my deck anymore 😂
231
396
3,971
696,947
Finally got around to doing it ! I’ve built a user intent-preservation evaluation pipeline, grounded in research learnings from my time at DeepMind. The underlying insight is simple: an AI system can produce an impressive output while failing the person using it. I translated the research into a framework covering four failure modes: user agency erosion, capability overshoot, user identity violation, and user intent misalignment. I then operationalised those into a scoring rubric and an LLM judge, and connected them into a working pipeline for batch evaluation of chat transcripts. The aim is to assess whether an AI system preserves users’ goals, constraints, control, and self-representation, not just whether its responses look good. Diving into AI UX success beyond task level scores. The next step is validating the judge against independent human assessments and testing what it catches beyond existing evaluations.
1
63
This is ideal to Run when ordinary success metrics stay healthy but interaction traces show the system is winning the task while the user is losing the thread. The clearest triggers are behavioral, not model-quality scores. Watch for repeated corrections that do not stick: the user restates a constraint, tone, audience, or “don’t do X,” and a later turn drops it. High edit distance after an accepted answer, frequent undo or regenerate on the same request, and sessions that end with the user rewriting the output themselves are stronger signals than a low star rating. So are abandoned flows after an apparently complete answer, short follow-ups such as “no, I meant…,” and users narrowing scope mid-task because the system expanded it. Product metrics that should prompt a sample are a rising gap between task-completion or thumbs-up rates and retention, reuse, or downstream acceptance of the artifact. Look also for more tool actions than the user asked for, memory or profile fields changing without an explicit update, and recommendations or drafts that converge on one style even when the user keeps supplying counterexamples. Support tickets and session notes that mention “it ignored me,” “it decided for me,” or “that’s not how I work” are direct reasons to score those transcripts. It is less urgent when failures are simple factual errors, refusals, or latency. Those already show up in standard evals. This can be used for the cases where the answer looks finished and the user still had to fight the system to keep their goal, constraints, control, or self-representation intact.
1
19
KP 🇺🇸🇦🇺 retweeted
i'm ready to book so many flights
271
481
16,027
398,043
It is easier to imagine the end of the world than the end of capitalism. It is easier to imagine the end of the world than the end of poverty and war. Ideas of an apocalypse seem to come easier than any utopia.
26
Insane
Introducing Higgsfield Ads Studio. The first ROAS-driven static ads content machine. > Does 10,000 hours worth of deep research of your brand > Monitors real-time trending formats > Crafts winning ads based on best industry practices Try free in Higgsfield ChatGPT extension.
51
KP 🇺🇸🇦🇺 retweeted
LinkedIn is becoming 90%+ AI posts: not just the comments but so many posts summarizing a well-written article using AI In all fairness, LinkedIn brought much of this on themselves: 12 months ago they added a button to write with AI; inviting AI writing (to juice post numbers?)
Using @pangram on LinkedIn is like seeing the AI matrix
155
34
1,060
69,645
KP 🇺🇸🇦🇺 retweeted
Using @pangram on LinkedIn is like seeing the AI matrix
111
51
921
163,070
KP 🇺🇸🇦🇺 retweeted
I've found that investors have a hard time investing in highly complex businesses because they're hard to understand. The ironic thing is that hugely correlates with the potential for building a moat.
66
47
682
53,139
KP 🇺🇸🇦🇺 retweeted
Just a reminder that tech week is every week in SF Go to events, get zero ROI, then rinse and repeat
6
3
69
2,360
KP 🇺🇸🇦🇺 retweeted
Million Dollar Homepage is reborn every decade in a different form.
TikTok girl selling ad space on her wall to pay for tuition not bad kid
7
26
571
45,385
KP 🇺🇸🇦🇺 retweeted
This one is for the mma junkies I fell in love with MMA when i found myself in an Irish pub December 2015 the night of the Aldo / Connor fight. There are big limitations with this approach. MMA rotates camera angles, specific models are less developed than the tennis ones, and grappling leads to constant occlusion This still seems very insightful as a fan who wanted to study what made mcgregor so special!
3
1
20
1,034
KP 🇺🇸🇦🇺 retweeted
Me and @BrainsAndTennis motion sequenced every Conor McGregor fight • His karate stance disappeared over the years • Fired pull/slip rear lefts 4x/3x more • His left was only avg speed! • Better cardio than he gets credited + can u guess the fight? nicodunks.github.io/Notoriou…
30
17
222
15,805
What’s stopping someone from Creating a fake hiring company, and running everything via agents, in order to obtain human data and work examples and information simply for research and data purposes?? I suspect a few are already doing this
1
5
84
The Healy generation kids are losing their minds rn
Shift Robotics launched the Moonwalkers Dusk this week. Powered shoes that strap over the ones you're already wearing, add force to each step, and get you to 7 mph without breaking into a run. $1,199, shipping late October. The headline is the gait model. Their S.E.N.S.A. system learns your stride inside about 10 steps, tells walking from running from jumping, estimates the slope you're on and adjusts assistance to match. Foot gestures switch modes so you're not pulling out a phone. 3.5 to 4.2 lb per foot depending on size, 90-minute charge, 6 to 7 mile range, 220 lb rider limit, IP54 — splash-resistant, not waterproof. But the line in the release that caught our eye at RobotAutoTrader is much further down. The straps, plastics and wheel rims come off without tools. That sounds like a footnote and it isn't. A consumer robot keeps its value when a second owner can replace the parts that wear out. The moment a cracked plastic or a worn wheel means a service centre, resale collapses — because a buyer is being asked to pay for a machine with unknown remaining life and no way to restore it. Wheels wear out. Batteries fade. Whether that's the end of a $1,199 machine or the beginning of its second life mostly comes down to whether you need a screwdriver.
2
293