Technology solving problems... and creating new ones

TechTalks retweeted
Politics, intrigue, and inter-lab rivalries are dealing a lot of damage to AI companies that depend on closed frontier models. While OpenAI might have scored a minor point against SpaceXAI by ending its patnership with Cursor, it has also set the precedent that makes it clear closed AI labs are not reliable. The correct strategy is to rely on open-weight models as the main engine of your applications.
OpenAI’s decision to sever ties with Cursor proves that relying on closed frontier models is an existential vulnerability for application-layer software. bdtechtalks.com/2026/09/04/o…
1
4
691
TechTalks retweeted
When I stumbled on the "Stealing Reasoning Traces from Proprietary LLM APIs" study, my first thought as a software engineer was, why did they overlook one of the key basic security practices of any software, which is to avoid global encryption keys? And if you want to hide the reasoning traces, why even send them to the client and not just store them on your own servers? It turns out that there are many tradeoffs and no perfect answers. On the one hand, AI companies want to prevent distillation (hence the encryption) while on the other, they want to be able to reuse conversations across different models (hence using a single encryption key) and minimize client data storage on their servers (hence sending the encrypted reasoning trace to clients). Basically, they wanted to have their cake and eat it too. In practice, they ended up messing up a lot of stuff: - Smaller, less safe models could be passed the conversation history and prompt-injected to reveal the reasoning trace - Client-side data filters failed to detect sensitive data stored in the reasoning trace, resulting in a lot of API keys, plain passwords and other sensitive data being leaked in public repositories - And in the end, competitors could easily bypass the anti-distillation guardrail
AI providers tried to hide the internal reasoning of their frontier models. A cryptographic oversight turned that protection into a massive enterprise data leak. bdtechtalks.com/2026/08/17/l…
4
6
733
TechTalks retweeted
AI security benchmarks mostly focus on models alone. But in practice, the model is deployed as part of a system, including the harness. A new study by @LassoSecurity shows how, all other things equal (including the model), the AI harness can have a huge impact on the security footprint of the AI system. This has direct implications for how you build applications as well as how you evaluate the security of your AI applications.
Developers often treat agent harnesses as neutral wiring, but new red-teaming research shows that your choice of harness can make or break your AI security. bdtechtalks.com/2026/08/11/a…
4
3
242
TechTalks retweeted
When your AI agent has thousands of skills/tools to choose from, stuffing all the skill details into the prompt is a non-starter. Passive skill retrieval is also not the solution. The key to managing hundreds of tools for your AI agent is to: 1) Decompose the task into atomic sub-tasks that each require one tool 2) Select the right tool for each sub-task and make sure tools are compatible 3) Create a DAG to execute the steps sequentially where required and in parallel where possible However, decomposing a task is easier said than done. For example, is an HTTP request one step or a sequence of ‘connect,’ ‘send the request,’ ‘receive the response,’ and ‘parse the response’? To solve this challenge, you need skill-aware decomposition (SAD), which uses the capabilities and boundaries of skills/tools to determine the granularity of each subtask. SkillWeaver and SAD are two techniques that can help scale AI agents to thousands of tools/skills.
Giving an LLM thousands of tools leads to noisy decisions. Learn how to optimize AI agent planning and tool routing without overwhelming the context window. bdtechtalks.com/2026/08/05/s…
1
4
13
970
TechTalks retweeted
There is a lot of promise and excitement in harness engineering for AI agents. The scaffolding that surrounds the LLM can have a huge impact on the speed, accuracy, and cost of your AI systems. However, harness engineering is a complex affair that requires a lot of manual work and experimentation. One solution is self-improving harnesses, an emerging field of research that enables AI agents to use the reasoning abilities of the backbone LLM to make modifications to its own scaffolding and improve its performance. The key to successful harness engineering is: 1) framing the problem in a way that makes it possible to optimize the harness in an AI loop, and 2) breaking the harness into distinct components that can be plugged into each other and composed like Lego bricks. I review two harness engineering frameworks, Self-Harness and HarnessX but there is a lot more to come. Keep an eye on this field.
With harness engineering becoming a main focus of AI engineering, new frameworks allow AI agents to write their own execution logic and optimize their performance. bdtechtalks.com/2026/07/13/a…
2
3
4
674
TechTalks retweeted
There have been exciting advances in self-improving AI agents that optimize their own scaffold/harnesses by leveraging the coding/reasoning capabilities of the underlying LLMs. Nvidia's ASPIRE framework applies the same principle to robot control models, using a continual learning loop to improve the robot control program. ASPIRE runs tests, examines execution traces, rewrites Python programs, and builds reusable skills that robots can employ as they face new tasks/objects in the real world.
ASPIRE and the new era of self-improving AI frameworks are drastically reducing token costs and deployment friction for real-world robotics applications. bdtechtalks.com/2026/07/06/n…
2
4
4
521
TechTalks retweeted
There's a lot of excitement around loop engineering for AI agents. But there is a clear difference between loop engineering and loopmaxxing. Loop engineering happens when you give your AI agent a clear, verifiable, and measurable goal, and let it iterate on its own output, analyze the feedback and adjust course to reach the goal. It usually results in increased productivity and visibility into how the AI agent works. Loopmaxxing happens when you don't put energy into designing the loop, give the AI agent a very vague and unverifiable goal, and hope that the underlying LLM will figure out how to do it. It usually results in the agent spiraling out of control and not achieving anything meaningful. Embrace the AI loops. Avoid the loopmaxxing.
The complete guide to the new loop engineering trend. Write powerful agentic loops while avoiding loopmaxxing. bdtechtalks.com/2026/06/22/a…
2
3
5
377
TechTalks retweeted
This is something I've been thinking a lot about lately: Chain-of-Thought (CoT) is the wrong direction to focus on. We spend billions on training LLMs to generate CoT tokens, hoping that they learn to reason like humans. In reality, there are multiple problems with CoT: - It is not a proper reflection of the human thinking/reasoning process (@rao2z has some interesting papers on this) - It slows down inference and forces the model to "reason" (if that is what we should call it) one token at a time - It raises the memory/compute costs of inference on long-horizon tasks What is the right direction? I think the model should reason in latent space and only use text tokens when it needs to communicate with humans or document its artifacts. Some interesting directions: - HRM and HRM-Text by @Sapient_Int reduce the costs of training and inference on reasoning tasks. - RecursiveMAS by researchers at UIUC shows how multi-agent systems can increase speed and reduce costs by passing information in latent space. Latent space reasoning is not without its challenges, namely the lack of interpretability. My suggestion is hybrid systems that use causal models to create a step-by-step plan for solving a problem, expressed in textual tokens, and recursive reasoning models that carry out those steps. Read the full article below.
Chain-of-Thought prompting is slow, expensive, and largely an illusion. The future of machine reasoning happens in latent space. bdtechtalks.com/2026/06/15/l…
6
5
18
1,556