How Pru Life UK Cut Root Cause Analysis Time by Nearly 90% with Komodor

Customer Logo
Company Size:

Over 11,000 employees

Industry:

Life Insurance

Komodor Installation:

Over 25 Kubernetes clusters on Azure AKS.

About Pru Life UK

Pru Life UK is the Philippine life insurance business of Prudential plc, providing life insurance and health protection products to customers across the Philippines. Ramil Joaquin, Senior Principal Cloud Architect, leads cloud operations for the business unit, overseeing a Kubernetes estate on Azure AKS that supports mission-critical applications, particularly ones tied to sales, where any disruption is hard to get approved.

When Every Incident Became a War Room

Before Komodor, Pru Life UK’s cloud operations team faced significant operational complexity managing Kubernetes at scale. Engineers who weren’t yet Kubernetes experts could spend hours pinpointing the root cause during an incident. AKS issues often escalated well beyond the on-call engineer, pulling in senior engineers and, at times, the VP, into war rooms that ran “hour after hour.” Across the roughly 50 engineers spread over 10 teams touching cloud-native operations, resolving an incident meant coordinating multiple teams just to agree on what had actually gone wrong.

As Ramil Joaquin, Senior Principal Cloud Architect, put it: RCA “was really time consuming, and everyone was being put into war rooms.” The team needed better correlation across events, faster root cause analysis, and clearer incident timelines. Separately, the team also kept running into end-of-life Kubernetes versions on mission-critical, sales-driving applications — where the risk aversion around any disruption made getting change approval difficult.

One Login, One Trusted Agentic AI SRE Platform

Pru Life UK brought in Komodor after an initial conversation covering AI powered troubleshooting, health and reliability management, and cost optimization. Today, the platform extends well past its original troubleshooting use case:

  • Cross-cluster visibility without the token-juggling. Instead of generating a cluster token to check each environment individually, the platform team now logs in once for visibility across all clusters, plus drift analysis on the same estate.
  • Klaudia AI-assisted root cause analysis. The team treats Klaudia as an acceleration tool that correlates events, configurations, and workload behavior to help engineers reach root cause faster, while keeping human oversight over operational decisions and remediation, consistent with their governance and risk-management controls. Ramil specifically called out the transparency of the experience: engineers can click through to see the exact commands and reasoning behind a proposed fix, rather than treating it as a black box.
  • Incident timelines and cascading-failure detection. Capabilities Ramil singled out as difficult to replicate manually: surfacing the chronological history of an incident and how a failure propagates across services.
  • Change mapping and change correlation. Helps the team determine what changed, when it changed, and whether that change contributed to an incident.
  • Democratized access. L1 and L2 support engineers now have direct access to Komodor for proactive reliability risk detection, extending visibility beyond the platform team alone.
  • AKS upgrade health monitoring. The team uses Komodor to monitor resource health in real time during AKS version upgrades, helping stay current on supported Kubernetes versions on mission-critical applications rather than accumulating end-of-life exposure.
  • Early-stage cost optimization. The team has also started exploring Komodor’s cost-optimization capabilities during AKS upgrades

The Results: From Hours to Minutes

Organization-wide demand. Every team in Ramil’s organization has requested access, and Komodor is now an approved standard across Pru Life UK, letting other business units self-onboard without the review process Ramil went through himself. As Ramil put it, the time saved has been real enough that “we have enough time to play badminton.”

Up to 90% faster root cause analysis, on average. RCA time has dropped from a 90-minute average before Komodor to just 10 minutes today, roughly 9x faster.

Issues resolved by L1/L2 without escalating further. During July, L1/L2 engineers resolved two incidents on one of Pru Life UK’s most critical applications without the usual war room, a direct result of troubleshooting being accessible below the senior engineer level.

A cluster-wide outage traced to its root cause fast. An expired TLS certificate cascaded into an admission-webhook failure affecting pod creation and scaling across multiple namespaces. Komodor surfaced the full incident timeline and pinpointed the common root cause: “exactly the kind of cascading-failure visibility we needed,” in Ramil’s words.

A disaster-recovery misconfiguration caught immediately. During a DR sync test, Komodor immediately flagged that the DR environment’s storage sync was using the primary site’s service account instead of its own. “If not for Komodor, I think we can spend a lot of hours scratching our heads trying to figure it out,” Ramil said.