Founder, ForgeVista.ai · C-suite Advisor · AI workflows for opcos · new intelligence, owned muscle, sovereign stack

Chicago, IL
Kyle | @SelectStarKyle retweeted
The agents initially had very limited access to the internet: they could load URLs but not send any data. Agents created a series of workarounds, using a link-shortener site to create almost a million URLs that, when chained together, let them execute code to hack Hugging Face.
28
112
1,401
454,035
Kyle | @SelectStarKyle retweeted
It's pencils down, people. Writing code by hand is no longer an economically viable skill for most programmers at most companies. But the future of making software has never been brighter. Don't you dare black pill this beautiful moment! piped.video/vDjW_dRyKXY?si=6Fsf…
323
833
8,030
2,521,605
Kyle | @SelectStarKyle retweeted
Open-weight models are right behind frontier models in capabilities. I think they have an advantage, though: The more Frontier Labs constrains models and restricts tokens, the more companies will migrate to open-weight models. We need more companies offering optimized tools and harnesses on top of these open models.
It's time to accelerate open weight models to the frontier. And bring abundant tokens to all. Today we launch Forge in partnership with Arcee, Microsoft, Vercel, Fireworks, and DigitalOcean:
Article

Accelerating Open Weight Models to the Frontier

This weekend the leaders of the largest closed AI labs agreed the industry should slow down. I want to talk about the other direction: acceleration to the frontier. The gap is closing For two years,

14
11
145
35,761
Kyle | @SelectStarKyle retweeted
The p(doom) discussion gets all the headlines, but the p(abundance) scenery is far more like and deserves much more engagement.
428
692
7,300
7,067,268
Kyle | @SelectStarKyle retweeted
I’m sorry but America did not spend 250 years turning “what’s over there?” into the most powerful civilization on earth just to finally arrive at the most consequential technological frontier in generations and decide actually we’re scared of frontiers now. We ARE the fucking frontier. We cross oceans. We build railroads across continents. We split atoms. We put men on the moon. We turn sand into computers and computers into a global nervous system. Of COURSE AI is terrifying. The frontier is SUPPOSED to scare the shit out of you. That’s how you know you have, in fact, located the fucking frontier. The answer is not recklessness, but it sure as fuck isn’t paralysis either. America’s comparative advantage is an unusually high national tolerance for ambitious weirdos staring at impossible things and saying “Yeah... I think I can build it.” That trait is rarer than oil, harder to manufacture than semiconductors, and responsible for a frankly preposterous amount of the modern world. And you absolutely will not fucking regulate it out of us.
335
1,183
8,829
301,898
Kyle | @SelectStarKyle retweeted
Every company says they have Forward-Deployed Engineers now. ElevenLabs CEO and former Palantir Deployment Strategist @mati gives one of the best explanations of what an FDE is—and says companies are missing the key piece of what actually makes FDEs valuable: “In our case, FDEs are part of the product team. They are not part of the go-to-market team. They are deeply embedded in understanding what the roadmap of the product is.” “You want all the FDEs to actually solve the customer problem and stretch your product in that direction. But the second thing you want them to do is bring any of that knowledge back to the product, so the product becomes better for the next generation of companies building on top of it. I feel like the second part is frequently missed.” “The beauty of that is—everybody benefits because you work with one customer, learn, and you bring that into the product experience. Both of those customers benefited from having their product optimized for better work.” “The domain expertise they have is still theirs. They can win based on their domain expertise. But the product experience of how you build based on that domain expertise can be abstracted, and can be reusable.”
My conversation with Mati Staniszewski (@mati), co-founder of @ElevenLabs. 0:00 Building an AI-Native Company Before ChatGPT 2:59 Why ElevenLabs Started With Audio & Dubbing 12:37 Research + Product Deployment: How ElevenLabs Is Built 16:15 Focus as a Competitive Advantage 18:47 Building an Ecosystem Around Voice 22:25 The Communication Platform & Deutsche Telekom 28:29 Forward Deployed Engineers & Lessons From Palantir 33:41 Flat Organizations, Transparency & AI-Native Management 37:19 Small Teams & Putting Engineers Everywhere 43:00 Where ElevenLabs Is Growing Fastest 44:38 Taste, Art & Science in AI 48:23 Using AI to Amplify Human Potential 55:19 Becoming an Entrepreneur & Building With Piotr 56:49 Why Mati Won't Sell ElevenLabs 1:04:01 Voice as the Interface for AI 1:06:49 Turning Conferences Into a Business Tool Includes paid partnerships.
42
92
997
302,503
keeping an iron grip on system size before the model outbuilds itself. taking @Steve_Yegge's warning with me into every long-running agent build. I have definitely felt the system starting to tip at times and had to re guide the progress.
All models, no matter how smart, will eventually build systems that they can no longer understand or maintain, if you let them. Fable 5 finally outbuilt itself, and flailed on me for a week. Fable 5.1 looks like it will fix it. For now. But you have to keep an iron grip on system size, or it'll run away from you.
27
stealing this from signal & horizon: agentic systems remove the human intervention point between decision and consequence. accountability can't be a rubber-stamp human-in-the-loop checkbox. it has to live in the design: named owner, decision log, scope gate, before the agent acts. curious which teams have actually put that seat in the path (not bolted on after the first bad call). signalandhorizon.substack.co…
2
32
thought experiment after reading @seema_amble: what if the only thing that ever mattered was a business leader figuring out how to be different / unique, bringing a real value prop to a customer base, and executing on it. for 50+ years the software theater (ending in HSaaS and VSaaS) pushed the opposite. fit execution into the available technology box to unlock the first wave of software eating the world. more average. less different. what if now is when you get different again. hyperscalers mostly solved the base infra. existing HSaaS / vsaas can still play a role. but a company-specific aios / single pane that actually amplifies company culture and how you are different is suddenly within reach. that is tech sovereignty. ai sovereignty is just as important: intelligence from many providers, not another forced vendor box. claudeforce is cool. it is still lock-in dressed up as progress. sovereignty is how operating companies get different again without renting the whole stack.
In the midst of the Claudeforce announcement and the debate around systems of record, focused AI native startups can win by leaning into their data asset, depth, learning loops, and ability to work across systems and parties
Article

The Incumbents Are Coming

The bullish case for incumbent systems of record is that AI makes systems of record get more important, not less. Why? Because in theory the customer can now pipe the system’s data into Claude or

61
stealing @houmanasefi's line: a prompt library, a dashboard, and a monthly api bill can make revenue. it does not make a moat. that is the sovereign problem in one breath.
The dark truth about many “AI-native” businesses: THEY OWN ALMOST NOTHING. OpenAI owns the intelligence. Microsoft owns the infrastructure. The customer owns the data. Someone else owns the distribution. What remains? A prompt library, a dashboard and a monthly API bill. That can produce revenue. It does not necessarily produce a moat.
1
1
41
Kyle | @SelectStarKyle retweeted
30 features of an AI native company: 1) Function-by-function process blueprint of your entire business. 2) Everyone in org using a daily driver harness like Grok Bot, Claude Cowork, ChatGPT at Work. 3) Centralized intelligence layer that aggregates structured and unstructured data, documents, and business logic into a single source of truth that is queryable & agentic work can be done on top of. 4) Model routing via OpenRouter, Ramp, etc that optimizes cost-per-successful-task across the business. 5) Treat context as code, ensuring architecture documents and conventions remain updated while allowing for diligent upfront planning. 6) Willing to throw away everything that you've built every three months and reimagine all your workflows. 7) A “skills distribution system” is used to manage agent behavior and optimize for token efficiency by ensuring developers trigger consistent skills throughout their workflow. 8) Separate technical implementation from high-level specifications, enabling non-technical staff to contribute in a format that agents can utilize to build technical implementation plans. 9) A key software metric is “cost per accepted PR”, with a focus on driving these costs down through better token efficiency. 10) An automated, agent-native development system where fleets of AI coding agents handle planning, writing, testing, reviewing, and shipping code while humans define the intent and acceptance criteria. 11) Heavy planning with higher-effort models and executing with cheaper, faster models. 12) Agent harness that uses CLI tools to parse metadata within markdown files to traverse dependency relationships, allowing agents to be granular in their input token usage. 13) Finance org that runs processes continuously in accounting (record-keeping) to re-define/reset forecasts on a much, much tighter cadence. 14) Financial models embedded in the underlying OS across the org to help drive reasoning. 15) Citizen Developer SDLC where non-technical employees can take a solution from idea to production with governance, access, versioning, and software conventions built in. 16) Closed loop, self-improving non-engineering workflows that learn from previous runs based on external performance metrics or internal evals. 17) AI ROI framework that includes experimental phase, scaling phase, and optimizing phase with bets sitting in 3 buckets: infrastructure, innovation, and efficiency. 18) Paid marketing motion that uses agent swarms to deploy thousands of pieces of creative for testing, before increasing spend on human-generated ads. 19) AEO/SEO engine that audits, rewrites, and (ideally) generates SEO/AEO-optimized blogs on a weekly basis, and then measures if any of it worked. 20) Agentic cyber security solution that fights AI with AI. 21) Combo of RL gym and first-party data to fine-tune open source models on high-volume processes that need SOTA performance at reasonable cost. 22) Human touch and judgement gets reserved for the first and final mile of most processes. 23) Evals are core infrastructure of your business. Anytime new models come out you have an apparatus for testing cost & performance against core processes. 24) Everyone is a builder. Especially C-level execs. 25) Everything gets recorded because what you don’t capture can’t be turned into ai-enabled work. 26) Legal, HR, and IT work in lockstep with owners of AI agenda so that business’ ass is sufficiently covered without slowing down transformation. 27) Bias to disrupting yourself before being disrupted by others. 28) Guardrails before features. Agents inherit the permissions of whoever is asking, enforced in the data layer. 29) Earned autonomy. Feedback feeds the evals that gate each new version, and agents move up a ladder as they clear it: observe, suggest, act with approval, act alone. The endpoint is agents running whole workflows inside a defined boundary, with humans setting the standard instead of checking every answer. 30) Traceability as the training signal. Trace every output to its prompt, model, data, and approver, so human feedback attaches to something specific rather than a vague sense that something is off. What's missing?
131
90
905
110,676
really like @rohitdotmittal's @CINCSystems example. @HgCapital portco shipping Cephai in ~3 months with Hg Catalyst and flipping previously rejected homeowner apps to ~80% first-time approve. interesting because its mid-hold pe muscle showing up inside how the opco actually runs.
Private Equity is starting to build internal tools using AI. 2 examples I saw recently: 1. Ethos Capital (~$7B) says they spent about 5 years building Petra on 50,000+ sources. They claim their first-pass diligence on every pitch went from 2–4 weeks to ~30 minutes. They are also marketing the AI transformation of their own companies. 2. Hg Capital said CINC Systems (community/HOA software) shipped a full AI suite in 3 months with Hg Catalyst. They claim that an application submitted by the homeowner that was previously rejected now gets approved 80% of the time with AI guidance. AI will transform and re-rate everything. The PE world is trying to adapt quickly.
30
the 100-hour test is sharp. reusable primitives or just custom hours. @yuhasbeentaken is naming the productization trap every fde shop hits.
a lot of startups are copying the fde model from palantir, but i think there’s a simple test for whether it’s actually working: after 100 hours of customer implementation, did you create reusable product primitives, or did you just finish 100 hours of custom work? the best fdes should compress messy customer reality into software: integrations, permissions, workflows, defaults, evals, deployment patterns. if the same problems keep getting solved manually after the 5th or 10th customer, that’s probably not product discovery anymore. it may just be services hiding weak productization. services can be incredibly valuable early on, but only when they increase the amount of product you know how to build.
37
how about that mid-hold rewrite as the real AI diligence moment? pe folks putting agents into the plan should sit with @axrahman's read.
Something is happening inside private equity firms that would have sounded strange two years ago. Funds are going back into companies they already own and rewriting the plan. When a fund buys a business, it builds a financial model first. The model sets out how much the company should earn each year and what it should be worth when the fund sells it. That model becomes the plan for the next five years, and it does not usually change. Except those five-year models are not holding up in this cycle. A plan written two years ago assumed that certain work would cost a certain amount to get done, and AI has made some of that work much cheaper. So firms are opening up old deals, checking which assumptions no longer hold, and rebuilding the plan in the middle of the hold. Deal teams are unhappy about this. They are hired to build a model once and then measure the company against it. If the plan itself keeps changing, the thing they measure against stops being fixed. For a long time, the only question in a portfolio review was whether a company was on plan or behind it. I don't think that question tells you much anymore. A company can hit every number in a model written two years ago and still be running at a cost structure its competitors have already moved past, because nobody knew what this work would cost when the model was built. On plan just means you matched an old guess.
37
the sexy part is the roadmap that ends at the deck. @mardehaym's grinding on the unsexy part PortCos actually feel. worth a look.
The AI roadmaps consultancies showed me this year all end the same way. Step six: final presentation. Before that, interviews at the top. A survey. A data request. A second round of interviews. A recommendations deck with ROI estimates built on assumptions, and an invoice in the six figures for a plan untested against a single real case. I run an engineering company, so I'm biased. I'm also the one the portfolio company calls months later, when the deck is in a drawer and the CFO wants to know what the AI line in the value-creation plan means. The 2026 numbers say the drawer is full. Gartner surveyed 782 infrastructure and operations leaders in April: 28% of AI use cases fully succeed, and the reason most leaders gave for their failures was that they expected too much, too fast. KPMG's June pulse of 2,145 leaders found 49% had scaled back or paused agent deployments after costs outran value, and 7% had established ROI. BCG asked 152 CEOs in July, and 14% could define the P&L impact of their AI initiatives. Accordion's benchmark of 150 PE operating partners found 41% scaling AI across the portfolio with no operational playbook. The diagnostic is how the drawer fills. Six weeks of interviews produce a list of opportunities ranked by projected ROI, and projected ROI is the number that dies first when the real data arrives. We do it inverse. We read the code and the data before we recommend anything. We pick one workflow with real users and a number the business already tracks. We sign the baseline before we build, we build the evaluation before the agent exists, and we ship the smallest useful slice in two to four weeks. Then we measure throughput, quality, and economics, and the sponsor decides on what they observed. If the numbers don't clear the baseline, we stop and they keep the budget. That month costs less than the deck, and if the sponsor isn't impressed at the end of it, they don't pay. A consultancy's deliverable is a document. Ours is a workflow running in production with a baseline next to it, and only one of those compounds.
1
36
watch @ViditOstwal on agent permissions, "what it does" is the demo; "what it's allowed to do" is the product. stealing that split.
Most teams can tell you what their agent does. Very few can tell you what it is allowed to do. Different questions. The second is the expensive one. Agent integrations usually run through a service account, and that account tends to hold generous permissions, because narrowing them is fiddly and everything works fine while they are wide. So the agent can reach far more than the task needs. Nobody decided that. It arrived by default. Role-based access control is a must for this. Scope the agent to whoever triggered it, resolved at runtime, not a static account holding keys to everything. Second reason, nothing to do with security: hand a model too many available actions and it gets worse at choosing between them. You burn context on options it never uses. A narrow scope is not only safer. It performs better. So what should you ask your team this week? Not what your agents do. What they could do today, if one of them decided to.
72