For an agent that runs for years, the hard memory problem is not remembering more. It is knowing when a fact stopped being true. A summary of old history compresses what happened. It cannot learn that a field on a government page moved. A fact should carry its source, expire when that source changes, and take down only what was built on it. @SidraMiconi tested this with AfterFire at the Long Horizon Agents Hackathon by tokens& (@tokensandai), at the AWS Builder Loft (@AWSstartups) in San Francisco. AfterFire is a disaster-recovery case agent. She moved a date on FEMA's page without telling it; it noticed and re-planned only what changed. Pre-registered race, 18 simulated months, 202 fictional households. Actions on a false belief: full history 47, summarizer 11, AfterFire 11. The tie is an honest null; a weekly recheck sets that floor. What differed: full history told 24 renter households the homeowners' date. AfterFire caught all 1,020 changes; planner cost $68.40 against $7,144.52 for full history. Summarizer cost was not reported. One simulated race is a data point, not a law. The FEMA dates are real; the households are not. @SidraMiconi · @AgentGlassInc - @nimble_search · @liquidai · @tinybird · @bfl_ai · @inaccisland · @ale_amenta · @yanivm13 · @ramin_m_h · @mlech26l · @vivmarquez · @jorgesancha · @bnevilleoneill · @robrombach · @andi_blatt · @pess_r · @OpenAIDevs · @awscloud · @AWS nitter.net/SidraMiconi/status/210…
At the AWS Builder Loft (@AWSstartups) in San Francisco, I built AfterFire: a disaster-recovery case agent whose facts expire when their source changes. For disaster case managers keeping a household's recovery straight across FEMA, the insurer, SBA, permits and contractors for years. Each fact carries a tripwire on the exact field it came from. I move the renters' date on FEMA's page without telling the agent; about half a minute later it notices on its own. Hundreds of facts retracted, only the affected parts re-planned, renters' texts drafted, not sent. Pre-registered race, 18 simulated months, 202 fictional households, three memory designs. Actions on a false belief: full history 47, summarizer 11, AfterFire 11. The tie is an honest null. Full history told 24 renter households the homeowners' date. AfterFire caught all 1,020 changes; planner cost $68.40 vs $7,144.52. Nimble (@nimble_search): live-web answers with citations; in the run on screen the planner got 42 tokens, not about 2,600 from the raw page. Liquid AI (@liquidai): LFM2.5 reads Maria's Spanish insurer letter on-device; nothing leaves the machine. Tinybird (@tinybird): append-only log; the case then vs now in milliseconds. Black Forest Labs (@bfl_ai): FLUX.2 edits only the moved line on her card; a second model reads it back. The FEMA dates are real. Maria is fictional. Code: github.com/Sidra/afterfire-h… YouTube: piped.video/watch?v=Re4AkXiW… Long Horizon Agents Hackathon, by tokens& (@tokensandai). - @inaccisland · @ale_amenta · @yanivm13 · @ramin_m_h · @mlech26l · @vivmarquez · @jorgesancha · @bnevilleoneill · @robrombach · @andi_blatt · @pess_r · @OpenAIDevs · @awscloud · @AWS . @AgentGlassInc
2
5
2,809
At the AWS Builder Loft (@AWSstartups) in San Francisco, Liquid AI (@liquidai) made its case for running agents at the edge, and the core of it is that inference on the device has no per-token bill. Tianshu Yu, an ML engineer who designs Liquid's vision-language model architectures, laid out the lineup: text models from 350M to about 2.6B parameters, a 3B vision-language model, and audio. On-device via llama.cpp, or vLLM and SGLang for high concurrency, or through OpenRouter. Four reasons for the edge: inference is free, latency is ultra-low, privacy (their healthcare partners keep data on their own servers), and it works without reliable Wi-Fi. Three hard parts: small enough for a phone, fast on CPU and NPU rather than GPU, and good at agentic tasks. His line: "nowadays we care more about the model doing things rather than being a chatbot." The 2.6B model was trained extensively for exactly that: computer-use agents, visual grounding, tool calling. @SidraMiconi · @AgentGlassInc - @tokensandai · @inaccisland · @ale_amenta · @ramin_m_h · @mlech26l · @vivmarquez · @awscloud · @AWS
3
1
5
134
At the AWS Builder Loft (@AWSstartups) in San Francisco, Tinybird (@tinybird) said it forked ClickHouse to make it schemaless. Send any raw JSON event, query it with SQL, never write a migration. It is called RawTree, and it launches in the coming weeks. Tinybird has run managed ClickHouse since 2019. Enzo Kajiya, who leads US sales, is testing a tagline: "If MongoDB and ClickHouse had a baby, it would look a lot like RawTree." Built agent-native from the start: an MCP server, skills, a CLI. OTel-native. Three use cases: observability, real-time analytics, and what he called sandbox and agent observability. An agent can query the data, and it can investigate its own work. Also auto-scaling, bring-your-own-cloud and bring-your-own-bucket. The hackers in this room get to build on it before launch. @SidraMiconi · @AgentGlassInc - @tokensandai · @inaccisland · @ale_amenta · @jorgesancha · @bnevilleoneill . @awscloud . @AWS
2
8
763
At @stripe: Mark Moriarty, Stripe (@MbyM): what it's like to be a Stripe. Chethan Shriyan and Hasan Tariq, AWS (@awscloud): AgentCore Payments — an agent hits a paid resource, gets an HTTP 402, and operates inside a budget enforced outside the model. David Mataciunas, M11 Labs (@DeividasMat · @m11labsai): selling to agents, verifying claims, and the cost of adding another agent to the loop. the model wasn't the difficult part of these demos. The difficult part was the money around it — who pays, who authorizes, who verifies... @SidraMiconi · @AgentGlassInc - @StripeDev · @patrickc . @stripe @Werner . @awscloud @MbyM @DeividasMat . @m11labsai
5
4
672
At @CapitalG, In Production, @OpenRouter's first quarterly showcase. Four talks nobody coordinated, one question under all four: what a model will and will not do when nothing is forcing it. Om Buddhdev, Olam Labs (@sensho · @olam_labs) : an arena where models play social games against each other and humans, every behaviour graded. The average human ranks around 14th, roughly Gemini Flash level. Gregor Zunic, Borwser Use (@gregpr07 · @browser_use) : less control, not more. The first demo took at least 25 tries. Buying something online is "one of the final bosses of web agents". Hai Ta, Userlens (@HaiTa · @userlens_hq): a cheaper model for ordinary clicking, more intelligence only after a bad experience. Demoed on CloseRouter, his OpenRouter clone, built because OpenRouter was not a customer yet, in front of OpenRouter's CEO. Alex Atallah (@alexatallah · @OpenRouter): the routing data. Content writing is dominated by Opus. Security audits go to open weight models, fewer refusals. His eval for every new model: "go clone Slack". @SidraMiconi · @AgentGlassInc - @OpenRouter · @dasha_shunina · @anastasiacrew · Alphabet Inc · @DemisHassabis · @SundarPichai
2
6
607
Same night, same stage, opposite advice. @dbreunig: a couple of installed MCP servers cost you 30-50k tokens before you press enter. Ship four of the ninety tools. @jpadamspdx: 500 tools in context every turn is "not tenable — expensive, and slow, and ineffective." @PrathmeshPatel_: "having too many tools in your MCP server, most major clients actually handle them fairly effectively." Fix the descriptions, not the count. Where I land: Prathmesh is describing where it breaks and Drew and Jeremy are describing what it costs, and both are right, which is the whole problem. A client that copes with 500 tools is still paying for 500 tools every turn. Nobody on stage had to hold both of those at once. Cutting all three into one piece, I did. nitter.net/SidraMiconi/status/210…
One MCP server, one tool, five agent frameworks. When the tool failed, three of them handed it to the model as an ordinary result, one raised a retry exception, one threw an error. Same tool, same server, same failure. This is the clip I'd send anyone who thinks MCP settled what an error is. It didn't, and it couldn't. Whether a tool call "failed" turned out to be a property of the framework, not the tool, and the model never gets told which framework it's living in. @ptdamiba : "a protocol success can contain a tool execution failure." That sentence belongs on page one of the spec.
1
4
109
How do we make agents actually work reliably in production? Evals, security, context, UI, tool routing, interoperability — a lot to unpack. MCP Community Connect SF was one of those days where the same question kept surfacing from completely different angles... @pamelafox @liamchampton @GlobAICommunity @SidraMiconi . @AgentGlassInc @hboelman @codewithsimon @pamelafox @liamchampton @digitarald @PrathmeshPatel_ @JiquanNgiam @dbreunig @sofiiiiiasz @cedricvidal @jpadamspdx @ptdamiba @TobinSouth @GlobAICommunity @neo4j @TryArcade
1
2
812
iPhone Duo in 60 seconds. Apple's first foldable: 7.6-inch display open, passport-sized closed, A20 Pro, from $1,999, pre-orders October 16. @SidraMiconi · @AgentGlassInc Chapters 0:00 iPhone Duo 0:04 Closed, it's the size of a passport 0:09 7.6-inch inner display 0:14 A precision hinge 0:18 5.4-inch outer display 0:22 Grade 5 titanium 0:26 Two apps, side by side 0:31 iOS 27, reimagined for the fold 0:35 StandBy, anywhere 0:41 48MP Fusion cameras 0:46 One battery on each side 0:50 Also new: iPhone 18 Pro 0:54 From $1,999
2
5
882
The demo broke on purpose, and that was the point. Builder After Hours: Durable Agents, Tuesday night at 101 California, San Francisco. @MelGoesTech of @temporalio killed a worker mid-ingestion, about 400 documents in flight. The workflow did not die. It waited, and when the worker came back it finished from where it stopped. @mattenoble of @KeycardAI switched off the app's access to MongoDB while the agent was running. The agent was denied and, in his words, "Of course, LLM is going to LLM." @NinaLopatina of @MongoDB cut a live system over from voyage-3.5 to voyage-4 embeddings by switching collections. Three talks, every Q&A, one reference architecture. The full event is in the reply below. — @SidraMiconi · @AgentGlassInc
2
6
1,033
The full event, on YouTube with chapters: all three talks and every Q&A, including the joint session where one question from the floor, how do all three of you respond to a rogue agent, is answered by Temporal, Keycard and MongoDB in turn. @MelGoesTech @temporalio @mattenoble @KeycardAI @NinaLopatina @MongoDB — @SidraMiconi · @AgentGlassInc piped.video/5fzWXZTDh1U
1
4
224
An AI that watches every session replay so a human does not have to. Patrick McConnell, Technical Account Manager at @posthog, on why: "if you've got hundreds of these, your eyes will fall out of their sockets." 1:10 - The case for never watching replays manually 2:18 - "are they getting pissed off and rage clicking?" 6:21 - "we record the DOM, capture video, and send it into Gemini" @SidraMiconi · Founder @AgentGlassInc
1
1
41
An agent that reads X every morning and DMs you the leads. @ishaansehgal, CEO at @omnaraai, opened the night with it: "an agent that every day at 10am gives us a digest of the five most important posts." 0:22 - Why the signal on X is so hard to find 0:41 - The 10am digest, defined 5:52 - A digest post becomes a DM'd lead @SidraMiconi · Founder @AgentGlassInc
1
4
55
"180 dials a day — by hour eight it gets hard to think of the next thing to say" David Vasquez, SDR at Outbound Sprint, showed a cold-call copilot built for exactly that stretch of the day: live objection handling while the call is still going. 0:22 - The hour-eight problem 4:06 - "the transcription came back thinking I was speaking Chinese" 9:29 - "using too much AI when you're selling can be absolutely brutal" @SidraMiconi · Founder @AgentGlassInc
1
74
"a 60% reply rate — 2.2 million in qualified pipeline from one event" @nehallagrawal, GTM Engineer at Invigilo AI, closed the night with the event-ROI brain: turning $30k of conference spend into a measurable pipeline number. She opened with "I was told not to prepare a slide, so now I'm wondering if I'm underprepared." 0:12 - The no-slide opener 0:42 - "wherever your revenue is coming from, my goal is to automate it and grow it" 4:04 - The reply rate and the pipeline number 7:55 - "I hate cold emails" @SidraMiconi · Founder @AgentGlassInc
1
1
44
Telling the model to "be brief" does not fix AI slop. @ToriSeidenstein, Co-founder at @tadata_team, on the bar a GTM agent has to clear: "when you hire an executive assistant you'll spend days calibrating them. From an AI agent you want it instantly." 0:49 - "we have all been there with AI slop, and that's not what we want" 2:06 - "our first approach was just telling the thing 'be brief' — that failed miserably" 4:07 - "we tested tons of models — the OpenAI lightweight models beat Grok" @SidraMiconi · Founder @AgentGlassInc
1
1
38
"a billion people on this new surface and very few people advertising in it" The surface is ChatGPT. @anfi_monaco, CEO at @Thrad_AI, walked through the arbitrage window on a brand-new ad surface — and what advertising inside it actually looks like. 0:58 - The billion-user gap 1:34 - "in the same way we have vibe coding, we have vibe marketing" 7:11 - "in two prompts you can build a company" @SidraMiconi · Founder @AgentGlassInc
1
1
39
The most interesting GTM demos right now are working agents, not slide decks. Eight ran live at AI Tinkerers San Francisco's GTM Engineering Demo Night, with @attio: an agent that reads X for leads, ads inside ChatGPT, an AI that watches every session replay, a copilot for hour nine of a dialling day. Thanks to the following presenters: 0:06 - @ishaansehgal: The X signal agent 11:08 - @anfi_monaco: Advertising inside ChatGPT 21:50 - The LinkedIn connection pipeline 28:05 - @ToriSeidenstein: No-slop GTM agents 39:28 - Patrick McConnell: The session replay scanner 47:05 - @yoeven: Selling to developers 55:20 - David Vasquez: The cold-call copilot 1:06:35 - @nehallagrawal: The event-ROI brain @JHeitzeb · Founder @AITinkerers (the SF team — @jacoblaes, @piersonmarks, @rjchint and Ian Butler.) @SidraMiconi · Founder @AgentGlassInc
1
10
3,978
Open any GitHub commit, add .patch to the end of the URL, and the author's email is right there. That trick is from @yoeven, Founder & CEO at @interfaze_ai. The demo: selling to developers — catch an engineer at the exact moment the problem is theirs. It opened on "today I'm going to teach you how to do GTM to people with broken legs." 0:09 - The broken-legs opener 1:01 - "engineers are the hardest people to sell to — if they already solved it, you lost the customer" 4:39 - "type .patch on the end of the URL — and there's his email" @SidraMiconi · Founder @AgentGlassInc
1
1
4
216
AI writes the code now. someone still has to read it. "review is one lane that never widens." @sachelik of Om Labs, at the @0G_labs event. teams ship more than ever. the bottleneck moved to QA. three clips in this thread. sound on. @SidraMiconi · @AgentGlassInc
1
3
3,135
the scoreboard is somebody elses exam. "question everything that your benchmarks tell you." @rabimba, at the @0G_labs event. be critical of your own benchmark, he says. then build your own. the setup for that line: models hand back outputs that confirm what you already believe. @SidraMiconi · @AgentGlassInc
1
3
126
a year of planned work, done in three months. his number. "this feels like crypto in 2018. it feels like mobile in 2010." @eshanchordia of Impulse AI, at the @0G_labs event. his read: keep experimenting or somebody else eats you. three voices, one thread. writing code got cheap. reading it, trusting it, and planning around it did not. @SidraMiconi · @AgentGlassInc
2
56