Introducing the Komodor Agentic Operations Platform

Itiel Shwartz. Introducing the Komodor Agentic Operations Platform
Itiel Shwartz
Komodor CTO & Co-Founder

In this webinar, Itiel introduces the Komodor Agentic Operations Platform, a robust system engineered to help organizations transition from localized AI tools to autonomous, production-grade cloud agents. It details the critical infrastructure gaps—such as permission management, guardrails, and continuous evaluation—that prevent standard coding agents from running safely at scale. To bridge this gap, Komodor Agentic Operations Platform offers pre-built solution models for immediate value and a flexible control plane that SREs and developers can use to securely build, deploy, and monitor their own agents.

TL:DR; Breakdown

  • Speaker: The session is presented by Itiel Shwartz, Co-founder and CTO of Komodor, with an introduction and Q&A moderation by Ilan.

  • Focus: The core focus is safely migrating AI workflows from local, single-agent laptop environments into fully managed, reliable cloud production environments.

  • Core Concepts: Critical requirements for cloud-based agents include Role-Based Access Control (RBAC), multi-layered guardrails, human-in-the-loop workflows, and continuous evaluation using “LLM-as-a-judge” methodologies.

  • Includes: The presentation includes a breakdown of KAOP’s architecture, specific pre-built solutions for AI SRE, Cost Optimization, and AI Software Operations, alongside a live demonstration of building and monitoring agentic workflows.

  • Wrap-up: Shwartz concludes that successful Agentic AI requires moving beyond “vibe evaluating” to implementing systemic feedback loops and rigorous security protocols, before inviting attendees to start a free trial.

Key Takeaways

  • Most organizations are currently stuck in the “local AI skills” phase, relying on developers to manually run and supervise agents like Claude Code on individual machines.

  • Migrating open-world coding agents directly to cloud production often fails because they lack essential infrastructure like automated RBAC, permission scoping, and execution guardrails.

  • Komodor identifies three primary use cases that drive the demand for agentic operations: AI SRE, Cost Optimization, and AI Software Operations.

  • Komodor Agentic Operations Platform provides a dual approach: out-of-the-box solution models for immediate day-zero value and a flexible SDK that allows developers to wrap and govern their existing custom agents.

  • Effective agent evaluation requires a “Swiss cheese” philosophy that layers datasets, manual validation, shadow testing, and LLM-as-a-judge systems to ensure continuous improvement over time.

Webinar Transcript 

Please note that the following text may have slight differences or mistranscription from the audio recording.

Ilan:

I will give a quick introduction and then let Itiel speak. Welcome everyone to Komodor. We are very excited to introduce the Komodor Agentic Operations Platform. This webinar will be recorded, so if you have to drop out, you will receive it later. Feel free to ask questions in the Q&A. With me today is Itiel, the co-founder and CTO of Komodor. He has led all our AI initiatives for more than three years. He will share the story of how we got to where we are today and what we have launched with the Komodor Agentic Operations Platform. Take it away, Itiel.

Itiel:

Cool. So like Ilan said the webinar is going to be recorded, feel free to ask like Q&A anytime you want. Like it says here like 55 minutes but I’m guessing it will be more closer to like between 20 to 30 minutes so don’t worry and I hope that I’m not going to take all of your time.

A bit about me, like very shortly like Ilan said over the last six years or like five six years I co-founded and was the CTO and still am of Komodor. We started as a company that was focused around Kubernetes troubleshooting graduated graduated it into AI SRE and now we are pleased to show the latest announcement or like the latest thing that we have created which is KAOP, Komodor Agentic Operation Platform and we are going to cover that in the upcoming slides. Other than that I really like the LLM space like creating agents, evaluating agents, trusting agents and that is pretty much everything I’ve been doing over the last three years before that a lot of like Python and backend and Kubernetes, things that are now becoming less important maybe with this new agentic area.

And a bit about Komodor like I said we started more from Kubernetes three years ago we did a switch to focus on AI SRE when ChatGPT was still quite new and young. And since then we were able over the last year to take it to the next level. We took everything that we have built, all of the infrastructure, all of the building blocks that we have built in order to create the best AI SRE in the industry and to allow our customers to build, run, manage, govern agents using the same infrastructure. So it means that what I’m going to show you now is how we build Claudia internally, Claudia was our own AI SRE, and we now give every customer the ability on one side to use or to get the best AI SRE out of the box given our solution model, but also to use the same building blocks that we used and perfected over the last couple of years to build, run, manage his own agents. Other than that we have a couple of quite known customers everything from like absolute and Dell and Cisco to booking affirm priceline open table and so on.

What’s the agenda for today? I’m going to talk a lot about the problem space like why did we do did this like switch basically why did we open everything that we have built over the last couple of years to allow our customers to build their own agent and we’re going to talk about why is it that hard to build production grade agents by yourself or without some guiding hand. Then we are going to talk about the main use cases that we saw in the industry and why we decided to tackle each one of them head on. And we’re going to talk a bit about the architecture of KAOP which is Komodor Agentic Operation Platform. I’m going to touch on guardrails and evaluation very quickly. I think it’s one of the most interesting topics to discuss and I think it’s going to shape how agents are going to be using production. And we’re going to end with a demo. Maybe I’ll like to switch to the demo even in the middle just so everyone won’t get bored. But that’s the agenda for today. Like I said feel free to ask questions at any time like Ilan. If there is something relevant just stop me and I’ll try to answer to the best of my knowledge.

Cool. So let’s talk about the problem. Like why like why are we all here? I think that one of the biggest things that we see right now in the industry is every company is going through pretty much the same phases which is going from manual work and basically trying to solve things by ourself to this phase where I think like 90% of the companies currently at local AI skills. We have skills for support, we have skills for triaging issues, we have skills to interact with Datadog. So we are doing a lot of things using code but most of it is still running as part of our laptop as part of our computer and it’s not really running in an automated fashion. Some companies are already here in the single agent area meaning they took at least like one of their skills and they are now trying to run it maybe on top of Kubernetes maybe Bedrock maybe Anthropic managed cloud. And on top of that we see companies trying to take this single agent and to make it as human-like as possible meaning to give him right permission to solve issues independently. And autonomous platform is where everyone wants to be where agents are taking the majority of the work from actual people. And I think that we’re going to see a huge change in the market as more and more of the task and the things people are doing will move from local AI skills into autonomous platform.

I think that like the problem or like one of the things that causes all of the problems or make all of those things even more acute is coding agents allow us as developers to move much faster. And that means that it’s now very easy to write code good code bad code much easier to write like bad code obviously. But we see that a lot of companies are now investing in how can I ship my code as fast as possible using like software factories and so on. But the other side of how can I use agents to automate everything that happens after my agents reaches production still is still very raw still very new. And this is the exact area where we want to give the best solution in the market like how do I take all of this velocity and speed that I already have in my coding agent and move it into the cloud.

So like I said I think like the biggest like revolution right of our time is like one year ago something like that with with Opus becoming as good as like a junior developer and cloud code becoming like a household name in every developer in every R&D in every company pretty much. So we see the same thing where basically we have some skill that works quite well and they work for us but in reality when we try to move them to the cloud we found that a lot of things are missing. Like if I’m working with cloud code usually he’s going to use my own credentials with my own permission he’s going to mimic everything that I’m doing. And I’m going to be the one triggering it. Once I take this cloud code from the local laptop into the cloud what’s going to happen is things are going to go south basically and we’re going to see a lot of questions that are starting to rise like like what do I actually need to do in order to take this local skill and move it into the cloud.

And I talk about this briefly like what are the differences and what are the problems that we see. And I think that like I have here like two or like three or the most interesting stuff that people don’t really talk enough about, which is like RBAC and permissions guardrails, how can I make sure my agent actually does things that make sense as well as evals. And how do I know that things are actually working as they should. When we’re working with cloud code a lot of the times he get things wrong he writes like a bad piece of code. But we are there to help him and maybe get mad at him but correct him. So I’m telling Claude hey build me a website the website is bad I’m telling Claude hey Claude the website is super bad like improve it the website is still bad. I’m telling hey please improve A B and C and then it improves it improves it. What happens in each one of those iteration is that I’m actually serving as a judge I’m telling Claude Claude you have done a bad job please reiterate please fix it. And once I move Claude into the cloud I’m losing this feedback loop I’m losing my ability to reiterate and to improve it and that is like a huge problem for agents running in the cloud.

Like I said one second I’ll open like the air conditioning here so I won’t die in front of you there. Yeah.

So like I said there are a lot of questions that people are already like asking themselves when they’re they want to move this like switch from the localhost into the cloud. I think those are the things that are a bit more in the mindset of people especially I think like all of the industry is mainly talking about those two areas. Harness and model. Like what should I use cloud code SDK langchain ADK pie agno so many different tools as well as models. Like obviously we have like an open weight models closed models Opus 5 5.5 just was just released yesterday I think and it already like feels deprecated in some manners. And I think like those are the two main things that people are really thinking about but in reality there are those set of considerations and then there’s the question of like I said RBAC evaluation cost and all of the things that happen once I move from my local laptop into the cloud.

So in reality what happens is you think that but that’s what you’re writing like prompt model tools maybe like skills as part of that prompt but that’s pretty much that. And in reality when you try to move it into production you see a lot of different things that you need to to take care of in order to really trust it and to make sure it’s like production grade and scale.

So I think that we we already like can all agree that moving skills into production agents is hard. And it’s not something trivial. And there’s a really big reason why currently most companies are still in the local AI phase local AI skills phase and not autonomous agents. When we are talking with customers and we talked in the last year with like hundreds of different customers and prospects we see three main use cases that are making companies obsessed which is AI SRE is like one thing cost optimization and AI software operation or AI SLDC. Those are the things that capture most of the attention currently in the market. We are trying to use agents in order to improve each one of those solutions. When I say AI SRE it’s also in the broader scope of things like support NOC and so on and when I’m talking about cost optimization it’s true both for like things like cloud cost but also for agent cost optimization cause we have a bunch of customers that their first agent cost them like 100K in the less in the first year even it almost provided like zero value. And AI software operation is where most of the hype is currently at things like code reviews, code factories coding agent sandbox agents and so on.

Okay so we covered like the use cases we talked about why is it that hard now I want to talk a bit about what we’ve built inside Komodor.

So obviously we tried to tackle the most interesting or the most painful use cases that we see for our customers and we have released like four different things as part of this new release. One is what we call the solution models. We have a dedicated out of the out of the box agents workflows screens capabilities for AI SRE cost optimization and AI software operation. So that means the moment you are onboarded on KAOP you are getting out of the box value and experience for each one of those three without spending one minute on writing agents. On the other side we also provide this platform a single platform that allows you to build run manage govern agents regardless of the type of agent. We try to make the platform as SRE platform dev NOC support as possible meaning if you are tackling one of those like use cases we probably got you covered in terms of how do we create the MCP connections the LLM gateway our ability to provide you evaluation and judges out of the box. So we build the platform for the operation folks to use and to build but it is like a platform that allows you to take your harness your cloud code and deploy it without needing to worry about the rest of the things. So we have again like three solution models as well as one single platform that everything is built on top of.

And yeah like I cover like this part of the solution models and the platform. I want to discuss like one area that was the focus for us in the last again like three years which is AI SRE. AI SRE is something that is super hyped right now we see dozens maybe hundreds of different companies trying to attack it from different angles. And we see cloud code when I say cloud code it’s also cursor and all of his friends are also trying to like tackle that specific problem like all of the APMs. In the end we saw that there are black box AI SRE solutions those solutions try to provide you with like a lot of value out of the box but usually fail because the world is super complicated is super complex and you and your team probably have a lot of domain knowledge around how do I solve problems how do I solve issues and so on. On the other side of the equation we have cloud code which is basically like a open a open word that allows to do everything that you want you can build everything you can do everything you can tweak everything but you are not getting anything out of the box. And that makes it extremely hard to rely on when it comes to building a reliable AI SRE that will run and investigate and remediate issues in production. We believe the the way to go is something in the middle meaning a platform that on one side allows you to bring your own knowledge agent harnesses into it but on the other side it also provide you with the ability to get a lot of the value out of the box meaning on day zero like two minutes after you install KAOP or log into KAOP you are already getting all of those agents out of the box incident pipeline incident flow. So there’s a great mix between build it yourself and black box operations.

Cool so we covered like this part. Now I’m going to talk a bit more around the technicalities. I hope I won’t bore you guys with that. And I try to make it brief and go to the demo cause I assume a lot of the people here are developers and it will be quite interesting to them to see a full live demo of the system and of the platform.

So let’s say that you already have an agent or even like a skill. What we do is we give you a very slim SDK that allows you to interact with the KAOP platform. Meaning you take your agent you install the SDK once you do that you are starting to get a lot of observability into the agent on one side but also a lot of capabilities out of the box like permission guardrails credentials and triggers. So you don’t need to do a lot of work in order to start with KAOP. All you need is an agent or even a skill you connect your cloud code you add this very slim SDK and this agent becomes part of the KAOP platform and gets a lot of the benefits that we’re going to cover in a second. How does it actually work how everything works together in the end of the day you have your agents you have the KAOP control plane and every time a request comes the control plane is routing it to the relevant agents and then it is routed to the relevant model or inference provider and allow you basically to have one single place to build to run to manage to guard to optimize your agent over time.

I don’t want to cover like too many things the only thing that I will say before like talking about the demo is AI SRE in particular and again this is where we bring a lot of our value is extremely hard and I think everyone that was a production engineer or SRE or developer knows that is because the world is super messy. Like in a lot of the demos of like using agents we see very simple use cases where it’s not that complicated or it’s quite static. If you are managing an environment with dozens, hundreds or thousands of different Kubernetes clusters where in each one of those clusters you might have hundreds or thousands of different pods it becomes extremely hard to manage and to guard all of it. And we know AI SRE is one of the biggest like frontiers for agent or for agentic operations and I think that by focusing on that we were again that was our core over the last couple of years we were able to build a very very strong foundation a very strong like building blocks that allows us to solve both the AI SRE problem but also problems that are a bit less complicated using the same mechanism the same guardrails and the same building blocks that we had in order to build the AI SRE system.

I won’t cover it. I won’t talk too much on like how can we simplify your AI SRE process mainly because I want to demo it to you and not having too much slides. I’m going to demo and then I think I’ll jump back to talking about guardrails and evaluation as those are quite interesting topic but I don’t want to bore everyone here.

So sorry give me one second do you want to say something Ilan while I’m switching demos?

Ilan:

I will reiterate something I don’t think Itiel touched upon but it’s it’s important for us to know I don’t know how many people noticed in the in one of the in the slides one of the important things is that the new Komodor Agentic Operations platform can can run anywhere so you can run the control plane on the Komodor managed cloud, on your own on your own cloud, on your own prem, air gapped, etc. So that is a significant difference and I think where we see a lot of customers also gaining value in utilizing agents so just reiterating that.

Itiel:

Okay cool that’s a good reiteration. Cool. So I’ll do the demo I think we have I’ll try to keep it like short like five to eight minutes and then I’ll jump talking about guardrails evaluation and we are done. So without further ado let’s start the demo and it’s like a live like demo environment so bear with me if we’ll see any kind of glitches and or problems.

So what we see here is the KAOP platform for everyone who knows Komodor already or is already like a happy Komodor customer. Everything like I said that we created is using the same building blocks but it is a new platform in terms of like new UI new capabilities and so on. What we see here is a fleet of agents I’m able to see all of my different agents where do they run what inference are they using and basically what’s the the status of each one of those agents. I’m going to show for a second how does the creation looks like and I’m going to do like a very quick demo of creation of agent using the platform. Important to say everything that I show you now is also available using our MCP. Meaning you don’t need to like go over all those click click click click my assumption is cloud code is going to do it better for you. We released I think so far 50 different skills in order to interact with KAOP meaning cloud code or cursor will work like that for any kind of feature or any kind of screen that you are able to see here.

Cool so we’re creating a new agent a research brief something to help us to like smart research on production product questions or suggestions. So I’m going to have the agent MD as the first thing that I’m adding then I need to choose the relevant provider and model. Like I said we provide out of the box inference model for our customers so that means that you are able to provide our customers with the ability to use everything that they want. Then I’m choosing the relevant skills that I want to use the relevant tools if I want any and the relevant agents that I want to connect to that specific agents. Relevant triggers how do I want this to start RBAC and permission who inside my company should be able to access that specific agent as well as relevant guardrails in terms of like budget and time so I won’t get bankrupt by using this agent. Evaluators is the LLM as a judge the thing that runs after the agent and allows me to grade it I don’t need to do it but we do recommend having judges to improve the agent. Memory is what should this agent remember and how can he reutilize those memories. And lastly is where do I want to deploy it again it can be in my cloud in like any other cloud provider and I can use any kind of helm chart and in the end the agent looks something like that and I can create the agent.

Cool that was like a very quick agent creation. I think we have a question maybe?

Ilan:

Yeah Chris asks how is the agent control and how do you make it deterministic?

Itiel:

Yeah so we don’t necessarily make the agent deterministic. The agent is controlled by the SDK so you can think about the SDK as a wrapper for your agent we call the agent a worker in KAOPs. And that means that every time we are going to trigger it to chat with it it’s going to go through this rapid. So you can use your own cloud code ADK pie whatever kind of harness and agent model that you want. And we are going to make sure that we wrap it with all of the relevant guardrails and permissions and RBAC and so on. You can make the agent deterministic but I think it does like defy the purpose a bit cause what we try to do is to take coding agent or LLM basically and to make sure that they work. Everything that I showed you now you can also say like screw that I don’t want that what I have is already a running agent and in that case the installation is going to be much simpler. You are simply going to do pip install or go I don’t know add to install the SDK and add it to your agent and that’s pretty much it so you don’t need to go over the UI you can do everything using your own agents your own docker files and so on. We are we give like docker files and helm images as like the basic recipe to deploy agents but in reality you can also deploy it using pretty much any kind of technology VMware sandbox on prem mainframe everywhere there is a compute we will be able to to work with.

Cool so I just showed like the different kind of creation agent now let’s see the support support triage agent which is one of my agents that is doing support. I’m able to see the number of runs how much did it cost me the latency the errors the text changes over time everything that interests me around the agent as well as what are the relevant connected integrations. And I can again like build run govern and improve my agent I think that the most interesting thing that there is is first of all maybe to look at one of the specific runs and to see the full transcript and to see everything that happened as part of this run and basically to better understand and analyze that. But on the other side we also have here like one of the guardrails like things that the company defined as a policies and how was this specific guardrail enforced during that specific run and I can also add it to the data set like if I want to say oh this is a great run I want to use it for further runs I’m also able to do that to give it some score this oh sorry it’s Hebrew this this is so great and I can do that and then next time I’ll be able to see it as part of my data set which allow me to evaluate improve and guard my agents from bad changes basically.

Cool so I just talk about run history a bit about the observability and revisions which is basically how did my agents change over time. And I want to talk for a second about the guardrails which like I said I’m able to limit both in terms of different building blocks guarding blocks basically as well as boring stuff like budget permissions and all of the things that you really need in order to make your agent run in production. Improve is the last part and when we are talking about improvement we’re mainly talking about data sets evals experiments and optimization. Data set is like the unit test if you want of the agent things that were captured and their status as part of the platform. Evals is using LLM as a judge to see the quality of the agents on different rubric. Experiment is allowing our customers to deploy two versions of the same agent and to see which one is better and run it as shadow mode or A B testing. And optimizations are us giving you guys recommendation around how to improve the agent what should you do mainly around like cost accuracy and time.

So I covered like I at least I hope that I covered like those part here around how do I build an agent and how can I keep on improving that specific agent.

Cool so we just watched how like a normal agent look like now I want to talk about the models the solution models and the first one that I want to focus on is the AI SRE. And when we are talking about AI SRE like I said it is built using the same building blocks that every customers of Komodor is getting but those are very specific to a specific use case in this case AI SRE. And what I want to show you now is how do I create what we call the workflow. Why do we need a workflow and what is a workflow? In order to achieve complicated stuff our feeling and our experience shows that creating a single agent is simply not enough. What you need is a bunch of different steps different agents working together to achieve the same goal and this is a workflow something that is predefined in a way something a bit deterministic combined with a lot of LLM in between so you are getting the best out of the two worlds. On one side you have this like high level blueprint of how things should run how things should behave but inside every part is actually an LLM which allow me to understand to to empower that using like non-deterministic rules and routes.

I’m going to show like a very quick example of here where I’m able to see all of the different alerts that arrived to the platform and I can see that this platform that is connected to Splunk observability cloud is currently not getting any kind of special treatment out of Splunk observability cloud. Meaning that alerts that arrive to the platform are simply not being handled by agents. What I want to do is to handle them using agents so I’m going to create a workflow in order to manage those specific agents use cases. And I’m going to do the create workflow and basically what I’m going to see to walk you through now is a full how do I create a workflow for managing incidents giving our recommendation and here everything is like predefined using our AI.

Cool so without further ado I’m going to first of all choose the trigger and I want to manage all of the alerts coming from Splunk observability a bit of like de duping and combining different alerts into a single incident. This is the most interesting part I’m going to build the investigation team that are going to help me to help solve issues for Splunk. And we have on one side investigation orchestrator you can think about him as like the quarterback or the incident commander he’s the one telling everyone how to solve issues how to solve problems and so on. With him I’m also choosing my team I’m choosing individual agents that will help me to solve the problem. We have here a combination of predefined agents by Komodor agents that we recommend out of the box as well as agents that the customer already created for that specific use cases. Now I’m going to hit next and I’m going to choose the remediation strategy for that specific use case and because most people don’t trust agents in production I’m going to say human in the loop so don’t do anything autonomous don’t try to solve the issue by yourself I’m going to choose like a validation strategy in order to explain what to do once the issue is resolved should I open a ServiceNow incident should I put it into Slack and so on. Again choosing the relevant pattern of security and who can access this specific workflow as well as should I override the relevant model should I have specific limits and spans for that specific workflow.

And in the end this is how the investigation team is going to look like. I know it’s quite scary but in reality this is how real production use cases of our customers looks like. Dozens of different agents trying to solve issues together as a single team for very complicated and cumbersome tasks. So I’m going to create this workflow which is great and now I want to show you how a workflow that already run looks like. So again I’m able to see it in terms of like accuracy the cost the latency and the error rate and I’m able to switch between to see how did it change over time. If I want to see how did the investigation looks like which I think is what interests most people this is an example of a single run of a workflow solving issues. I have here like the timeline I have what happened why did it happen and the full investigation and remediation that happen for that specific agent. And I can even see the full investigation and basically this is how a real investigation looks like if in the previous graph we showed like the blueprint of the workflow this is how an actual workflow looks like. From our experience a traditional investigation require between six to eight different iteration and between 20 to 30 different tool calls before you are able to reach the root cause of the incident or of the issue.

So we had here like an example of investigation the relevant guardrails and so on of that specific sorry investigation. But if I’ll go back to the workflow I’m able again to see the agents that are composed of that the observability metrics evaluation quality lab again everything that we saw for agent is also applicable for the workflow itself as well as opportunities experiments and so on.

Cool like we are already like 30 minutes and I did promise to keep this webinar a bit short so I will if we have questions Ilan just let me know and if not I’ll talk for a second maybe on like evaluation guard rails and we’ll wrap it up.

Ilan:

Yeah just just one quick cause I answered the questions live but if you want to add anything on top of that so we did have a question if you can add Komodor Agentic Operations as a Slack teammate or a Teams teammate. I answered that we do and I believe we’re also like dog fooding it internally and actually using it like that but if you want to add anything to that Itiel.

Itiel:

I think that’s one of the most useful use cases, Slack, Teams, or WebEx as well as connecting it to your cloud code codex and so on. My assumption is most of the usage won’t be via the UI. Our goal in KAOPs in to integrate with your existing lifecycle your own existing flow your own existing tooling and to empower you where you are not forcing you to go to the UI. Yeah.

Okay that’s it cool. So I will jump very fast like again I don’t want to dwell too too much… So I’m going to show this thing here.

Cool. So I would talk for a second around the guardrail I think agent runtime and how do you protect it is like one of the biggest problems we identify five main like vectors that are most interesting it starts from where your user is getting the data then all the way to MCP, LLM providers and in each one of those area there might be a problem there might be a risk. What we are providing our customer is five and here you can see like all all of them different gates that allow you to make sure that your users are not and your users might be developers or support or NOC or each one of those will be able to operate the agent without worrying that he is going to delete the database or do some pod deletion or AWS destroy the account. Or on the other side to get bad like responses that are inaccurate and to follow up on top of them.

Last thing is evaluation I think that is like the biggest problem of our time in terms of agents like they do so many things how can I make sure that they are actually sensible basically and how can I make sure that I can trust my agents over time even that I’m changing it and the world itself is changing. What we came up with is like again three years of experience is knowing that in the end it’s more like a Swiss cheese kind of philosophy very similar to what Anthropic is also saying. You need different layers where every layer help you to catch or to to util or to guard on a different aspect we have like the data set and benchmark and manual validation and LLM is a judge and shadow testing and A B testing you need all of those different strategies in order to make sure that your agents are behaving in the way that they should over time. In KAOP what we do is we provide you with all of those out of the box meaning like you don’t need to do anything and you are already getting those and we allow you to have the platform that we with that allows you to build on top of meaning create your own LLM is a judge do your own benchmarks create your own data set so again on one side we bring you a lot of the value out of the box but on the other hand the ability to tweak and tinker as much as you see fit.

I think that the most like interesting area that we are and again I’m a really big believer is the continual learning loop the ability of the system itself to keep on improving over time. And we see that in the end you don’t really care about agents or models or harnesses you care about outcomes. And what we’re trying to do is to create this autonomous feedback loop that will allow your agents your workflow your use cases to keep on improving even where you are not there even where you are not next to the keyboard. And this is like something that I think we’re going to change how the industry is thinking about agents over time and I’m proud to say that it’s something that we are already supporting the platform allowing you to build and to keep on improving your agents over time.

Okay so yeah like 35 minutes has passed I think I covered most of the things that I was supposed to cover Ilan unless I forgot anything so with that last round of questions and if not we’ll wrap it up.

Yeah let’s see if there is let’s wait like if there is another question I will ask you this Itiel sort of like put you on the spot you mentioned you I know you’ve hundreds of calls like in the last like six months or so with customers and prospects and you’ve been like knee deep in this for like three years what would you say is like the biggest like one single thing that people overlook when they’re starting to work with like agentic AI is there is there like one thing that people like that gets overlooked by people that you comes to the top of your head?

I think that like between like security to evaluation I don’t know one of the two. I’m not sure which one is more crucial but I see people running agents in production and when we ask them how do you know that they work they usually like start telling us every ten oh that’s a good question I don’t really know but I hope it will work. And on the other side there’s the how can you make sure that it is not deleting your production and a lot of the times we see people are like yeah we are having our fingers crossed and we hope for the best so I think like those are the two main areas.

Ilan:

So from vibe testing vibe coding to vibe evaluating…

Itiel:
Yeah yes yes like I think that’s pretty much what we see across the industry if we don’t have any questions I’ll say that I am available over like LinkedIn and email and if you are looking for a single platform to empower your team or for a solution already ready for like AI SRE cost or developer operation like feel free to to ping us or to like start a free trial right and like start using the platform.

So thanks a lot thank you Ilan.

Ilan:

Thank you everyone and like Itiel said yeah you’re absolutely welcome to go on the website and register for a free trial or contact us or just reply to the email everyone everything will be received. And let you all go have a good night good evening good morning wherever you are and see you at the next webinar.

Frequently Asked Questions (FAQs)

Can the Komodor Agentic Operations Platform (KAOP) be deployed securely on internal infrastructure?

Yes, KAOP is designed to run anywhere, including the Komodor Managed Cloud, your own private cloud, on-premises servers, or fully air-gapped environments.

Do developers have to use the KAOP UI to build and deploy their agents?

No, developers can skip the UI by using Komodor’s slim SDK. By simply wrapping their existing code or agents with this SDK, they instantly gain access to the platform’s observability, guardrails, and role-based access control.

How does Komodor prevent autonomous agents from making destructive changes in production?

Komodor implements a layered “Swiss cheese” security model. This includes strict agent guardrails, Role-Based Access Control (RBAC), LLM-as-a-judge evaluation, and support for “human-in-the-loop” remediation to ensure autonomous agents cannot execute critical system changes without approval.