Komodor spent years building an AI SRE platform before the category had a name. With the launch of the Komodor Agentic Operations Platform, it’s opening that engine up so enterprises can build, run, govern and optimize their own agents in production. Following the launch, co-founder and CEO Ben Ofiri sat down to talk about why now is the right time for agentic operations, what breaks between prototype and production, and where operations will head next.
Why did you decide to build this now?
Interviewer: What is the Agentic Operations Platform, and why did you decide to build it now?
Ben: Our new platform is aimed at medium to large enterprises that want to build, run, and operate their own agentic workflows to support different use cases around SRE and operations, such as incident response and cost optimization, and other software operations tasks they want to automate.
What we saw in the market over the last 12 months is that there’s a huge demand for enterprises to build their own agentic operations. Everyone saw what happened on the coding side with coding agents, and now companies are trying to implement similar automation and methodologies in DevOps, production, and SRE practices.
The problem is that, unlike on the development side, these agents face different challenges. The scale is enormous — think about large enterprises with thousands of clusters and tens of thousands of virtual machines. Second, we saw a lot of challenges around guardrails and governance: how do you make sure these agents are really doing what they should, that there’s no scope creep, no hallucinations, and no actions that would be very difficult for the business to recover from? Lastly, how do you ensure these agents are actually doing their job properly — how do you measure their efficiency, and know whether they found the root cause or the right trade-off between cost and reliability on a specific AI workload or GPU cluster?
So there’s a very large demand for organizations to build their own agents, but also a lot of unresolved challenges around governing, optimizing, and visualizing those agents at scale. From everything we did over the last couple of years around AI SRE — using Klaudia, our own agentic engine that we built internally — we realized we’d already built a lot of the key components that could allow other enterprises to build their own cloud-scale agents. So twelve months ago, we decided to take all the expertise, tools, and methodologies we’d developed internally and make them part of a larger platform, so any enterprise could build, run, optimize, and govern their own agents with the same efficiency and quality we achieved with Klaudia — but for different use cases and needs.
Where does this fit in the AI SRE market?
Interviewer: Komodor is already known as an AI SRE platform, and as a company solving these problems before the term “AI SRE” was coined. What trend do you see in the industry, and how does the new product tie into that?
Ben: First, we’re seeing the AI SRE market grow significantly — analysts like Gartner are writing about it, and the latest reports suggest that within a couple of years, around 90% of enterprises will use an AI SRE solution in their production systems. So on one hand, everyone understands that automation and AI need to be part of SRE and production practice, to keep up with what GenAI is doing on the developer side.
On the other hand, we’re seeing that the vast majority of in-house initiatives to automate SRE with agentic workflows are actually failing. I think the reason is that it’s easy to spin up one agent, or write one skill, or even several — but it’s very hard to run a whole fleet of agents on your production systems 24/7 with the precision and efficiency that’s required. That gap, between the need to automate AI SRE and the ability to actually build and manage a fleet of agents running autonomously, isn’t addressed in the market right now, and we think we’re well positioned to solve it, given our experience building Klaudia and working with Fortune 500 and Global 2000 organizations on incident response, cost optimization, and more.
So we don’t see this as a pivot from what we did before — it’s the next phase in our evolution, from a black box to an open platform that every organization can use to build their own agents.
What breaks between laptop and production?
Interviewer: Can you talk about the gap companies face moving an agent from a laptop prototype into production?
Ben: It’s similar to the transformation that happened around 7 to 10 years ago with microservices. Everyone thought microservices were a contained thing — you’d just spin up a new one and each team would own its own — but what we actually created were very large distributed systems with tens of thousands of interconnected components, which brought new challenges around security, visibility, efficiency, and optimization. I think something similar is happening now with agents, but probably at 100 times the scale.
Organizations are trying to leverage GenAI and foundation models, and the first thing they do is take a prototype that worked in a few cases and try to turn it into an agent that lives in production. At that point, most customers think: I have a skill, some prompts, and a model — that’s it, I can productionize this into a real agent that runs 24/7. But what they’re neglecting is, first, governance: how do you make sure the agent is doing what it’s supposed to, that there’s no scope creep? If an agent is meant to do PR reviews, how do you make sure it’s not reading information it shouldn’t from the database, or taking actions like a rollback that it shouldn’t take? You need to understand how to delegate but stay in control — there’s no silver bullet, some things you want to block statically, others only dynamically, and you need an engine that lets you customize and enforce your own policies at scale.
The second thing is visibility and optimization. These agents are, by definition, non-deterministic, so you can’t just write tests and assume they’re doing their job properly. You need different methodologies, including evals, to understand whether the agents are actually working — because you might have tested 5, 10, even 100 cases, but now the agent is taking thousands of actions or decisions every day. How do you know it’s not hallucinating, that precision is high enough, and, most importantly, that it’s improving over time rather than degrading? That requires stitching together different tools to build a dataset, run evaluations and experiments, close the feedback loop, and make changes to prompts, models, and tools. That takes expertise, time, and a lot of bandwidth from your team.
The third thing, assuming you’ve handled governance and visibility, is making sure you’re driving positive ROI. Everyone is automating these SRE and production workflows to save time, improve efficiency, MTTR, and uptime — but tokens are expensive. We hear from customers that they have a system that works reasonably well, but it can cost millions of dollars a year. So you also need to think about optimization, not just performance, but cost: choosing the cheapest model, making fewer calls, caching results, and so on. A lot of companies don’t think about this when they start the journey, until a couple of months in, when they realize the year’s budget was exceeded in month one, and then they understand they need different practices to achieve positive ROI on their agentic initiatives.
So those are the main challenges we hear from customers, and our platform was built to address exactly those three: governance, visibility and optimization, and ROI.
Build, buy, or…?
Interviewer: Customers today seem to be choosing between an off-the-shelf AI SRE product, a vendor-provided one, or building it themselves with coding agents. What are the pros and cons of each, and how does Komodor address that?
Ben: That’s the tension we hear from most prospects and customers. On one hand, you have all these great-looking startups that promise their AI SRE just works — you install it, MTTR goes down, tickets get resolved automatically, costs go down. The promise is huge. On the other hand, you have bare-metal agents or generic frameworks you can use to try to build your own agents for your specific business needs.
I see significant disadvantages in both. With black-box AI SRE, the demo looks great, it solves problems, and so on, but when it meets real environments and real production needs, it breaks very fast. We built Klaudia, our own AI SRE, over more than three years to get high precision, high recall, and high consistency, and even after all that work, out of the box it supports only specific use cases at the precision we wanted: Kubernetes troubleshooting, GPUs, AI/ML, and so on. Organizations’ systems are complex and need a lot of customization to handle edge cases, and black-box AI SRE tools offer a system that’s supposed to work end to end, but when it breaks or you want to customize it, it’s very hard, because they already have their own approach and systems that are hard to influence or leverage.
On the other hand, generic agent frameworks force you to build all the troubleshooting, incident response, and cost optimization expertise yourself: building different agents and workflows, creating the integrations to monitoring solutions and cloud providers yourself, orchestrating everything together, managing conflicts and findings, and building your own evals and datasets to improve over time. That requires a lot of expertise, time, money, and bandwidth.
What we see is that most organizations don’t want to just trust a black box, but also don’t want to be forced to build everything, visibility, optimization, governance, from scratch. Our approach lets you enjoy both worlds. On one hand, we offer more than 50 battle-tested SRE agents that work out of the box for complex use cases, from troubleshooting an AWS problem to GPU cost optimization. On the other hand, we offer a complete, extendable platform that lets you build any agent you want with our own SDK, or import your existing agents using any coding agent, harness, or agent runtime you already use. After helping you build or import those agents, our platform lets you choose where to run them: your clusters, our cloud, or your cloud, and then optimize their performance and cost with our eval lab, and set the right guardrails and governance with our governance engine.
We believe in this blended approach, and we think it’s what makes customers choose Komodor as their agentic operations platform.
AI problem or SRE problem?
Interviewer: How does Komodor’s past experience solving complex production and SRE problems play into this? How much of this is an AI problem versus an SRE or DevOps problem?
Ben: It’s a healthy combination of both. On one hand, foundation models themselves are getting smarter and more sophisticated, and are now able to handle very complex, multi-step tasks. This wasn’t the case two years ago, or even a year ago, but with the new models released around nine months ago, we found they can do root-cause analysis for very complex cases that might require 20 different iterations and calls. So you already have this amazing engine that can produce great results and solve complex problems.
The main challenge customers have, if we use a metaphor, is treating these agents as the brain cells: they’re very smart and getting smarter, but how do you actually bake that into your enterprise and production system? What’s missing is something like a nervous system and a body that can take all those theoretical capabilities from the LLM providers and turn them into real, concrete, actionable results and findings in your production system.
Combining those two things is probably the hardest challenge right now. To do that, you need to understand not just how LLMs and agents work, but also production: site reliability engineering, what a DevOps or platform engineer’s role looks like, the constraints, and what they’re trying to achieve, and derive the best agentic solutions and workflows from there. For example, it would be naive to assume a large, highly regulated enterprise will let even the best model take write actions on its production systems directly: maybe it’s not allowed, maybe they need auditing, maybe they’re required to have a human in the loop for every action. You need to understand the domain of site reliability engineering, what’s allowed, what our audience of DevOps, SRE, developers, and operations teams actually do and want to achieve, and also how to leverage the advancements in LLMs and AI to automate as much as possible.
Our team is composed of a lot of people with hands-on LLM and agentic expertise, but at least half the team comes from the SRE, DevOps, and platform engineering domain. I think combining those two domains is what makes our platform unique and lets it address the real pain points our audience has.
What does operations look like in three years?
Interviewer: If everything goes the way you think it will, what does operations look like three years from now?
Ben: If you’d asked me a couple of years ago, I’d have said not much would change dramatically. But given how the pace of advancement is accelerating, I think operations will look significantly different than it does today. I think the best analogy is self-driving cars: in twenty years, most driving will probably be autonomous, and ten years ago nothing was. Right now we’re in between — some cars have some autonomous assistance, some are fully autonomous, but only on specific roads and in specific cities. I think we can all feel intuitively that in a couple of years, a lot of the technical problems and regulation will be resolved, and the path to fully autonomous driving will open up.
I think what’s happening in autonomous operations is similar. Up until a couple of years ago, 99% of operations was done manually and required a lot of expertise and time. Think about how companies fixed incidents, or did cost or reliability optimization: most of it was derived manually, requiring a lot of people, headcount, and expertise. I think in the next two to three years, 80% of those processes will be automated using AI — some fully, end to end, some only partially — but overall, 80% of the processes will be automated. For incident response specifically, I’m confident detection, investigation, and triage will be automated, and in some cases remediation too, though some cases will still need a human in the loop if confidence isn’t high enough.
I think we can expect a very interesting path for autonomy in operations and DevOps practice, and we’re excited to be part of that.
See the platform in action. Watch Co-Founder & CTO Itiel Shwartz walk through the Komodor Agentic Operations Platform live, from guardrails and evaluation to building an incident-response workflow. Watch the session now – no registration required.
Ready to see it on your own environment? Book a demo →
