Web3 Research / AI fintech / Sports and more

Bimbadao
A robot copied a human move. That footage used to stay locked inside companies. Person picks → moves → places. Shift records it. Robot replays it. @openroboto turns real work into robot training data on Bittensor. SN80 rewards the footage. $SN80 live on Base + Robinhood Chain
3
1
7
131
A robot copied a human move. That kind of footage used to stay inside companies. In the clip, a person picks up a part, moves it, and sets it down. Shift records the gesture from a first-person view. The robot replays it. That is the scarce input in robotics. Not another model. Real hands, real objects, a real sequence. Most of it never leaves the lab that filmed it. @openroboto is opening that supply on Bittensor. SN80 rewards accepted footage. Companies pay for the tasks they need. After the exclusivity window, eligible data can open for research. $SN80 is already live on Base and Robinhood Chain. Human work in. Robot training data out. $TAO $SN80
SN80 OpenRoboto is breaking one of robotics’ biggest walls: data locked inside closed companies and single ecosystems. Shift turns real-world human work into robot training data, while @openroboto is distributing that network across Base and Robinhood Chain. Physical AI, opened up. bittensor:native robinhood:0x6f63d869011f95274498023b4abfc00b30c34378
7
2
39
934
Bimba retweeted
Google just dropped an absolute nuke on the AI leaderboards Anthropic held a complete monopoly on this market for months They were sitting comfortably at 80% to win the October cutoff Then Google unleashed Gemini 4 Argon The Polymarket chart violently ripped from 10% straight to 59% overnight This isn't a minor patch, it is a generational leap that is breaking the Text Arena: > blind test win rates are completely skewed > Claude is getting systematically cleared out in head-to-head battles > the strict Style Control filter explicitly favors Argon's new architecture Retail is still hesitant, leaving the contract mispriced at 59¢ polymarket.com/event/which-c… With the October 31st snapshot rapidly approaching, it is structurally impossible for Anthropic to train and deploy a counter-measure in time Buying Google right now isn't a prediction It is just claiming a guaranteed 41% yield on a mathematically locked race
5
2
55
4,160
Sonnet 5.5 + Jev, solar breath we keep paying models to redraw space nobody watches 16 seconds. 8 planets out of 8 still tracked. one coronal loop, LOOP-01, grows from 0.32 R☉ to 0.52 R☉. corona 0.378 → 0.760. wind 0.284 → 0.705. modes: observe → corona → wind → system. on top of that, one line: the sun rises even if you are not waiting. two ways to make the same clip: > send Sonnet every frame, every planet, every particle > Jev holds orbits, corona and wind - Sonnet only paints the sun breathing same picture. different bill. full render → you pay for Neptune, which was already tracked Jev + Sonnet → you only pay for what is actually changing why this makes money: > 16 seconds is a reel, not a film > the HUD already looks like a product, not a prompt > one loop, one track, one hook > people stay for the numbers, not for “wow, a sun” how to sell it: > Jev keeps the state: 8/8, corona, wind, loop > Sonnet keeps the light > you put one sentence on top > the post is not “I generated space.” the post is “here is the log” observation record saved at 14 seconds. that is the product. if your agent is staring at the sun, let Jev read the system once. let Sonnet breathe. 8/8 · LOOP-01 · corona open 100% · a new day starts anyway.
5
2
36
659
Bimba retweeted
i put an AI that can't code in charge of two that can I gave Jev, GPT-6.1 Sol and Claude Sonnet 5.5 one sentence: “add team accounts to this 118,000-line app without breaking billing for existing users” Jev didn't touch the code. it chose the next move every time the work branched: → GPT-6.1 Sol mapped how accounts, invitations and billing fit together → Sonnet 5.5 changed the files and wrote the tests → Jev sent unclear decisions back to GPT-6.1 Sol and routine work back to Sonnet 5.5 → tests ran after each finished part, with risky choices left for me the moment that sold me: Sonnet 5.5 found two ways to give an invited user access. one also gave that user access to billing Jev stopped the next edit and brought GPT-6.1 Sol in. it caught the permission path, Sonnet 5.5 fixed it, and the new test failed on the old code but passed on the patch across the run: 472 Jev decisions → 31 calls to GPT-6.1 Sol → 27 files changed → 186 tests passed 3 hours 26 minutes. $24.60 in API calls. the Jev decisions were about two cents of that I reviewed the billing and permission choices before anything merged GPT-6.1 Sol didn't spend the night making routine edits. Sonnet 5.5 didn't have to guess at the risky parts. Jev kept handing the next move to the right one one instruction went in. a working feature, tests and a reviewable diff came out full playbook below
10
7
39
1,407
Sonnet 5.5 + Jev in an arctic desert we've been paying Sonnet to watch the same white field 16 times Jev got the empty ground, two tumbleweeds and one live track: T-01, Ø 1.08 m, wind already writing on the sand. they tested two methods on the same 16 seconds: > Sonnet on every grain, every old wake, every dead segment > Jev holds wind, path and decay - Sonnet only rolls the body that is still moving the picture didn’t change. the work did. full field, every frame → you keep redrawing SEG-07 while it dies Jev + Sonnet → T-01 only · 3.12 → 6.08 → 5.00 m/s 16 s · drift → gust → trace → fade wind 7.4 m/s 323° → gust 10.5 → 8.5 m/s 306° path 121 m → 193 m trace 72.8% → 39.7% SEG-07 age 8.2 s → 23.2 s · open flat the reason is almost stupidly simple: each extra frame on a dead wake is the same document read again. Jev reads the field once. Sonnet stays on contact. if you're building this into a field stack: > put wind, path and grain into one Jev state > define the live tracks up front > batch drift, gust and fade > let Sonnet take the uncertain roll > send contact-lost cases to a human contact dropped 0.21 s and came back. the line was the same. only the old segment kept fading. if your agent is staring at an arctic desert, let Jev read the wind once. let Sonnet move the ball. T-01 · ridge crossing · the wind writes. the sand forgets.
9
4
41
817
i think i found Sonnet 5.5’s missing piece I put a fogged night forest in front of Jev. Sonnet 5.5 only had to finish the three bats that actually hung Jev held 18 tracks the whole way: range, heading, altitude, wing phase, and whether a roost slot was still empty 18 airborne → approach → landing → roost 16.0 s · 60 fps · tree T-03 · slots 0/3 → 3/3 dolly in 15.3 m → 5.7 m vis 61 m → 99 m airborne 18 → 15 roosting 0 → 3 for the same 18 bodies: → paint every flap as a hero shot → you drown in fog and false contacts → Jev keeps the flock map, flight traces and event log → Sonnet only draws the three that brake, spread, grip and invert about one tree, three landings, no wasted focus the workflow: > Jev cuts the sky down to inbound tracks > Sonnet reads the approach and holds the landing > tests are the feet on bark > a human watches the last three hang Sonnet still did the landing. it just didn’t spend the whole night sorting bats that were never going to roost B-07 · B-12 · B-03 BRAKE → WINGS SPREAD → CONTACT → FEET GRIP → ROOST INVERTED → SWAY DAMPED
12
2
42
1,035
Bimba retweeted
OpenAI claimed their AI solved the Navier-Stokes problem a few weeks ago Retail immediately rushed to Polymarket to bet on what gets solved next They are currently pricing the Hodge Conjecture at 28¢ They think they are betting on the future of mathematics But if you actually read the fine print on the market rules: > no verification from the Clay Math Institute is required > it only needs a public announcement > from an eligible lab like OpenAI, Google, or Meta Retail is betting that a neural network will genuinely crack another century-old problem by 2027 Smart money is just betting on which PR department drops the next press release to pump their valuation It’s not a math market It’s a corporate marketing market You are trading tech PR disguised as pure science
4
1
34
933
Bimba retweeted
i didn't expect the cheaper Claude to get this close to Opus Claude Sonnet 5.5 is out. on a coding test built from real Cursor sessions, its best score sits about two points behind Opus 5.5, with input and output tokens priced at half Opus's rates the bill: > $2 / 1M input tokens > $10 / 1M output tokens > $0.20 / 1M cached reads two published comparisons of Sonnet 5.5 against Sonnet 5: document work → 2.4x faster → 12% fewer total tokens → more accurate results coding, both at high effort → score 10 points higher → about 15x cheaper per task another early tester ran 2,441 finance tasks: → Sonnet 5: around 497,000 tokens per answer → Sonnet 5.5: around 121,000 higher scores with about 76% fewer tokens per answer I'd start with the work that keeps coming back: bug fixes, interfaces, spreadsheets and slides the setup: 1. give it one clear job 2. include the relevant files or screenshots 3. define what the result must pass 4. run the checks and inspect the output keep Opus 5.5 for complex jobs that need sustained judgment pick one smaller task and test Sonnet 5.5 using the setup in the guide below ↓
12
12
84
5,955
SHE JUMPED OUT OF A BALLOON ONTO A TRAMPOLINE IN THE OCEAN fifteen seconds. one woman. one basket. one floating circle with a red X. the clip contains almost nothing: → “let’s go,” then she leaves → no explanation of the platform → no safety brief, no setup, no why → she hits the X, bounces, climbs back in while the basket stares the missing noun is the whole post. viewers have to finish it: who put that trampoline there, who thought this was the plan, what happens if she misses. nobody can be wrong. every comment supplies the sentence the video refused to write. a like is cheap. inventing the logistics of a mid-air ocean trampoline is work. platforms weight the work. the jump stops the scroll. the unanswered why makes them type.
16
3
85
3,846
THE MONKEY SHOVED A BANANA ICE CUBE INTO A VOLCANO ten seconds. one animal. one block of ice. an open door over lava. the clip contains almost nothing: → no cut until the dump → no speech, no look at camera → ice goes in, fruit comes out → nothing labelled, no reason given the object itself is the implied antecedent. viewers have to finish the sentence the footage started: why the cube, why bananas, why the volcano, why the monkey decides. nobody can be wrong. every comment supplies the missing noun. a like is cheap. inventing the motive is work. platforms weight the work. the footage stops the scroll. the unanswered why makes them type. nothing rational happened. people will still explain it.
6
1
51
3,708
Bimba retweeted
i put Jev in front of 24,000 customer messages before Fable 5.1 saw a single one refund requests, late deliveries, damaged orders, duplicate charges. Jev checked four things for each message: > what happened > how urgent it was > whether the policy covered it > whether it needed a person 24,000 messages → 96,000 decisions → 1,200 cases sent to Fable Fable read the difficult 5%. The other 22,800 followed approved rules, with code checking the order and policy before any response went out I ran the same queue through both setups: → Fable on every message → 6h 42m → $187.20 → Jev first, Fable on the shortlist → 1h 19m → $24.90 about 5x faster and 7.5x cheaper across the full run I also checked 300 routing decisions against the policy. 294 were right; I corrected the other 6 before sending anything the handoff: Jev sorts → Fable handles hard cases → code checks → a person approves refunds Fable still does the difficult work. It just doesn't read 22,800 routine messages to find it set up Jev's first pass with this guide ↓
13
13
101
5,415
Bimba retweeted
Smart wallet on Polymarket just made over $8K trading weather markets He put money on NYC getting between 5 and 6 inches of rain Entered at just 37.9¢ and is currently up over 150% on the position His total net profit is over $8.3K And the most interesting part? This wasn’t a one-time trade He systematically enters rain markets across cities like Seattle and Seoul And maintains a perfectly smooth PnL curve across all his trades On just $130K of volume, he's already made 195 predictions I think this is only the beginning Definitely a trader worth keeping an eye on
11
3
34
1,102
This isn’t real Nations League footage. That Kylian Mbappé dictator clip was fully generated with Picsart. France without their emperor tonight. Belgium as the last country still standing. King Baudouin turned into a throne room with flags, gold, and a balcony speech. It looks more expensive than the actual broadcast from Brussels. France 45%. Belgium 30%. Draw 27%. $1.66M already in. Kick-off in minutes. People already can’t tell AI from a real match. The future is here. Most people just haven’t noticed yet.
2
2
24
714
This isn’t real NFL footage. Those stadiums were fully generated with Picsart. Ravens as a gothic cathedral. Buccaneers with a pirate ship and cannons. Steelers as a working steel mill. It looks more expensive than an actual game. 12M+ views in days. People already can’t tell AI from a real broadcast. The future is here. Most people just haven’t noticed yet.
6
42
2,145
Bimba retweeted
i didn't expect people to build this much with Jev already 1,000 emails sorted for roughly $0.03. 1,000 reviews checked in 4.6 seconds. fewer calls to Claude. 8 open-source repos, starting with the ones you can use every day: 1. sort your Gmail → jevmail sort Gmail into needs reply, updates, promos, sales and spam. each message gets an urgency score. the author reports: 1,000 emails → about 1 minute → roughly $0.03 read-only Gmail access. nothing gets sent or deleted. github.com/fazlerocks/jevmai… 2. send uncertain emails to Claude → claude-x-jev Jev sorts emails, leads or comments. Claude reads the uncertain items and writes. the author's inbox: 216 emails → 38 for Claude to review in the separate sorting comparison: → Jev: about $0.01 → Opus 5.5: $4.73 github.com/charlesdove977/cl… 3. let Jev handle browser clicks → firefox-jev-mcp Opus sets the browser goal. Jev handles the clicks. in a Wikipedia navigation test: → Claude turns: 12 to 4 → Claude tokens: 88,500 to 20,800 → total time: 23.4 to 14.7 seconds github.com/kitoutou999/firef… 4. turn app reviews into priorities → jev-column-race turn 1,000 app reviews into sentiment, topics, bug reports and scores for how likely users are to leave. → Jev: 4.6 seconds / $0.023 → Gemini 3.8 Flash: 18.8 seconds / $0.158 4.1x faster. about 7x cheaper in the recorded comparison. github.com/goodrahstar/jev-c… 5. find code by asking a question → every ask your codebase a question like "which functions ignore errors?" the author measured: 1,302 functions → 3.7 seconds → $0.018 you get ranked matches with files and lines to inspect. github.com/sufianetaouil/eve… 6. load only relevant skill descriptions → jev-skill-gate Claude doesn't need every installed skill description in every session. one Rust example with 217 skills: estimated description tokens → 12,750 to 3,185 75% smaller → roughly $0.00074 for the Jev call the skills remain available to invoke manually. github.com/ShivamPansuriya/j… 7. check commands before they run → hermes-jev-approvals Jev checks commands against your rules and returns approve, deny or ask a person. in the author's 156-command benchmark: → review time: 619 to 63.2 seconds → human prompts: 42 to 10 a separate study against GPT-5.4 mini found a smaller 1.24x reviewer-time speedup. github.com/anpicasso/hermes-… 8. put useful search passages first → jev-rerank-server put relevant search passages first before your model reads them. 25 search queries → 500 relevance checks 16.9 seconds → about $0.016 at listed API rates in that sample, the top-10 ranking score improved from 0.616 to 0.718. github.com/gbesse/jev-rerank… all eight are open source. the numbers come from the authors' tests and examples. API usage is billed separately. pick a repo → get your Jev API key with this guide → follow the repo's setup instructions ↓
14
9
54
4,085
Bimba retweeted
🚨GOOGLE USED HER AI TO BUILD ITS HARDWARE. AND HER STARTUP IS VALUED AT A F#CKING $4 BILLION At Stanford, she spent an hour breaking down the feedback loop that lets AI agents improve their own work: 10:20 - Stanford's "infinite monkey" trick for finding an answer that actually works 22:25 - the moment an agent tests its own code and tries a different approach 44:44 - why a customer support agent doesn't need to know every answer More steps won't make an agent smarter if it can't tell which ones worked. My article below takes that idea into a business workflow built with Jev + Claude Code.
12
4
40
1,185
Bimba retweeted
EVERYONE IS WATCHING HAALAND TODAY. I AM WATCHING 1-1. On @bagel_win I have two bets today, and both are about goals. Germany vs Greece: a combo. Germany to win + both teams to score. 68% vs 13%. I do not buy a clean sheet. Klopp needs his first win, so this game will open up. Norway vs Portugal: both teams to score again. The market is already at $761k and the sides are almost even. Haaland needs one moment. Two matches. One logic. Who is with the goals today, and who is still on the favorite?
19
2
39
4,411