Founder @ Randoli • Building AI-powered observability OpenTelemetry | OpenCost | KubeCon Speaker | Former Red Hat | 🇨🇦

Toronto
What would it take for you to use an AI agent to monitor your production environment? When I do demos and ask that question, I typically get two very different reactions. On one side, I hear: "I'm not going to let some AI agent access our production environments." What about permissions? Compliance requirements? How do I know the agent can be controlled when it can combine tools in ways we haven't thought of and potentially cause issues? How do we tackle the non-deterministic nature of AI when production environments require consistency and predictability? The other reaction is: "Why can't my team just put together a bunch of agents and tools and build these workflows ourselves?" There are a lot of MCP servers out there. Teams can combine them and build different workflows. It's the age-old build vs. buy argument. Both are very valid questions worth digging into. Let's start with the first. Security and compliance are non-negotiable for most organizations. Any solution that uses AI agents to monitor production has to address these concerns. At a minimum, you need to be able to: 📌 Explicitly configure which agents and tools can be used, and what permissions they have. 📌 Use pre-approved instruction sets (aka Runbooks) 📌 Trace what the agent has done, including actions taken on behalf of a human user. With the right guardrails in place, you can incrementally enable AI agents as you build more trust and experience. For the second question, I think we can take a page from platform engineering. The goal is to empower teams while still enforcing standards and structure. You can absolutely put together your own set of agents and tools. But it's easy to underestimate the effort required to build a proper framework around them that addresses governance, security, compliance, auditability and ongoing maintenance. Those are things you can't compromise on if you want to allow agents to monitor and eventually take actions in your production environments. I believe the use of AI agents to monitor production is inevitable. At some point, most engineering organizations will have to decide how much access they're willing to give them. And when they do, the hard part won't be connecting an LLM to a collection of tools. It will be building the guardrails, governance and trust required to let agents operate safely in production. #Observability #AgenticAI #SiteReliabilityEngineering
20
8
525
This is the mindset we need to have when leveraging AI. Otherwise you are not taking advantage of it. As a founder it can help you with planning, researching, product design, marketing, SDR, financial projections. Before I would need very skilled people for researching & planning and multiple folks to execute especially the grunt work. But now I could iterate faster and execute the same with a much smaller team.
If you're stuck at "AI is just another tool", you're making a category error. This is not "same computer, but better". It's far more akin to hiring a bunch of extremely intelligent and capable coworkers with whom you might occasionally disagree.
1
28
Getting a production alert on a Saturday morning is not ideal. Thankfully it's a minor issue. Site Reliability Engineering 😔
1
50
During the @open_cost community meeting we discussed about AI generated code contributions and how we want to handle it. On one hand you want to encourage participation. There are tons of things that can benefit from extra hands or agents ;). But at the same time we have to ensure the core is thoughtfully crafted. Weather we like it or not AI generated code will be the primary mode of contribution in the not so distant future. We all have to adapt to the new reality. Love to hear how other projects at @CloudNativeFdn are handling this. #OpenSource #Kubernetes
64
Data sovereignty is important in Observability too. Your traces and logs can contain sensitive informaiton. Where your telemetry goes matter, especially when you want to get AI to reason about them. #DataSoverignty #Observability
44
When we started building an autonomous SRE Agent, we thought giving it highly distilled, high-value signals would result in better reasoning. We were wrong. The more underlying context we gave it, the better it became at reconstructing what happened and reasoning toward a root cause. It makes sense when you think about how an experienced SRE investigates an incident. Metrics lead to traces. Traces lead to logs. Logs lead to deployments. Kubernetes events and infrastructure state help reconstruct the timeline. Remove too much of that context and you make the investigation harder. That led us to an interesting contradiction: AI wants context. Traditional observability economics encourage us to remove context. I wrote about what this could mean for the architecture of AI-driven observability. Article in the first comment.
1
38
This looks very promising. We are already experimenting with it on @openshift for our @randoli_inc autonomous SRE Agent that uses OpenShift MCP among other tools.
FYI OpenShell was submitted as a sandbox application to CNCF and was accepted a couple of days ago! #OpenShell #CNCF
56
Spot nodes gives you free chaos testing.
2
29
Some days I crave comfort food. The Ethiopian food platter is one of them. My wife was also craving the same so we ended driving into Toronto to have Ethiopian food tonight. While eating, it reminded me of the similarities between the platter and Sri Lankan rice & curry. We typically eat rice with a bunch of curries. Usually a mix of vegetables curries, lentil curry and a meat or fish curry. If you swap the Injera for rice and the spice mix to the Sri Lanka curry powder there's virtually no difference. Even the texture do the curries feels similar.
116
Rajith Attapattu retweeted
Replying to @shubh6200
Kind reminder. Not everyone is at Uber scale. The current database you use is likely more than enough. Besides being a 10 yr old article these kind of articles are a good read for stroking your intellectual curiosity. It's fascinating to learn how these planet scale architectures work. But real wisdom is knowing that these learnings are rarely applicable for most of what you do day to day. Ditto for what you hear at conferences.
1
2
12
1,996
I love it when a customer wants to collaborate and be more than a consumer of your product. The demo turned into a brainstorming session on how the Autonomous SRE Agent framework can be extended to tackle a wider set of problems spanning ops, customer support and engineering. I walked away from the call with ideas and a plan to collab. Great way to end the week.
What if we treated cost as another observability signal? Most organizations already have pretty good visibility into their cloud spend. AWS, Azure and GCP provide detailed billing data. FinOps teams have dashboards showing spend by account, service, business unit and sometimes application. But there is a difference between knowing what you spent and understanding why your cost changed. For Kubernetes, knowing that an AWS account spent $80,000 on compute last month doesn't tell an engineer very much. You need to be able to move through the infrastructure: Cloud Account → Cluster → Namespace → Workload → Pod → Container And then connect that infrastructure back to the organization: Team → Application → Product → Cost Center Now imagine Kubernetes spend increases 25% over two weeks. Knowing that the increase came from three clusters is useful. Knowing that most of it came from two workloads is much more useful. Knowing that one of those workloads doubled its CPU requests after a deployment gives an engineer somewhere to start investigating. This is where I think cost starts moving beyond reporting and becomes another form of operational telemetry. Applications aren't static. We deploy new versions. Traffic patterns change. Resource requirements change. Someone changes a request or limit. A workload starts scaling differently. Performance improves or degrades. So why do we treat cost as something we look at separately, often days or weeks later? I think cost metrics should sit alongside the other signals we already use to understand an application. Latency went down. What happened to cost? We increased CPU requests. Did performance actually improve? Traffic increased 20%. Did infrastructure cost increase 20%, 5% or 50%? A deployment improved p95 latency by 15%, but doubled the cost of running the service. Was that the tradeoff we intended to make? That's the feedback loop I'm interested in. Not optimizing cost in isolation. And definitely not reducing cost at the expense of performance or reliability. It's understanding cost, performance and reliability together as the system changes. If observability is supposed to help us understand the behavior of our systems, cost should be part of that picture. #CostManagement #Kubernetes #Observability
83
Kubernetes tip: Use PriorityClass for your critical pods to protect them from getting evicted when there's a resource crunch. Not every service in your cluster have the same priority nor do you want to create some sort of complex scheme by defining priority for everything. Just set it for the critical ones.
2
4
183
What if we treated cost as another observability signal? Most organizations already have pretty good visibility into their cloud spend. AWS, Azure and GCP provide detailed billing data. FinOps teams have dashboards showing spend by account, service, business unit and sometimes application. But there is a difference between knowing what you spent and understanding why your cost changed. For Kubernetes, knowing that an AWS account spent $80,000 on compute last month doesn't tell an engineer very much. You need to be able to move through the infrastructure: Cloud Account → Cluster → Namespace → Workload → Pod → Container And then connect that infrastructure back to the organization: Team → Application → Product → Cost Center Now imagine Kubernetes spend increases 25% over two weeks. Knowing that the increase came from three clusters is useful. Knowing that most of it came from two workloads is much more useful. Knowing that one of those workloads doubled its CPU requests after a deployment gives an engineer somewhere to start investigating. This is where I think cost starts moving beyond reporting and becomes another form of operational telemetry. Applications aren't static. We deploy new versions. Traffic patterns change. Resource requirements change. Someone changes a request or limit. A workload starts scaling differently. Performance improves or degrades. So why do we treat cost as something we look at separately, often days or weeks later? I think cost metrics should sit alongside the other signals we already use to understand an application. Latency went down. What happened to cost? We increased CPU requests. Did performance actually improve? Traffic increased 20%. Did infrastructure cost increase 20%, 5% or 50%? A deployment improved p95 latency by 15%, but doubled the cost of running the service. Was that the tradeoff we intended to make? That's the feedback loop I'm interested in. Not optimizing cost in isolation. And definitely not reducing cost at the expense of performance or reliability. It's understanding cost, performance and reliability together as the system changes. If observability is supposed to help us understand the behavior of our systems, cost should be part of that picture. #CostManagement #Kubernetes #Observability
1
164
A timely reminder that people make the difference and your team matters. Its not about tokenmaxxing or running a bunch of agents nonstop, but motivating and enabling your team to be creative and focus on producing the best experience possible for your customers.
In the obsession with AI, too many founders, leaders, managers seem to have forgotten that an engineer with motivation + drive + focus + AI >> hundreds of AI agents without any of those The best products and companies are built by people and teams like that, and will continue to be so Ask yourself, again, why the AI labs hire the best of the best (some of the most driven people in the industry; founders, CTOs etc) if people aren't making the differrence when everyone has access to similar tools already
1
150
When we built our SRE agent, the biggest bottleneck wasn't their intelligence, it's that we force them to use heavy, slow language models for basic decisions. Using an LLM to generate paragraphs of text just to figure out "Should I run this tool?" or "Is this route safe?" is slow and expensive. TypeSafe AI’s Jev seems very interesting here. It skips text generation entirely to act as a fast, intuitive "System One" brain for agents. It processes inputs and instantly outputs strict, typed choices and calibrated probabilities. By offloading operational routing and guardrails to a dedicated decision model, agents could become faster, cheaper and likely more reliable. Will post more after experimenting further. #AgenticAI #SiteReliabilityEngineering
4
2
184
We are about to drop another release. We've settled on a bi-monthly release cycle than big releases which has really helped us maintain stability while continuing to innovate. What allows us to do this is our own product. We show change markers within the main metrics which allows us to spot regressions quickly. #Observability #SoftwareEngineering
1
48
Uptime monitors helps you measure your API endpoints against your SLOs . You can now setup Uptime monitors in the @randoli_inc platform and connect it to an AI driven incident management workflow. You can specify Runbooks to be executed by our SRE Agent to investigate and potentially remediate issues supporting your SRE team. #Observability #Monitoring #Agentic
1
65
Is it just me or has the older claude models somehow got dummer in the last few weeks. (maybe resources being diverted away to newer models ?) I'm seeing more mistakes and subpar output.
1
1
70