1. Home
  2. Case studies
  3. Kuda
Case study · Cloud engineering

24/7 site reliability engineering for a digital bank

Observability, SLOs and round-the-clock on-call support for a mobile-first bank with millions of customers.

Client KudaIndustry Digital bankingLocation Lagos, Nigeria & London, UKDuration Ongoing since 2023
5 minalert response time
62%fewer customer-facing incidents
24/7engineer coverage

The client

Kudadigital banking, Lagos, Nigeria & London, UK.

The challenge

A bank cannot have a maintenance window. Kuda's product team was growing quickly, but reliability still depended on developers noticing problems in dashboards and fixing them out of hours. Alerts were noisy, incidents were resolved ad hoc and there was no shared definition of what "healthy" meant for each service.

What we did

We introduced a site reliability practice and staffed it around the clock together with Kuda's engineers.

  • Service level objectives for the customer-facing journeys: sign-up, transfers, card payments, statements
  • Consolidated observability with Prometheus, Grafana and the Elastic stack; alert rules rewritten around SLO burn rates
  • Opsgenie on-call rotations with escalation policies and runbooks for every alert
  • Blameless post-incident reviews and a weekly reliability report to leadership
  • Backup and cross-region recovery drills every quarter

Results

Customer-facing incidents fell by 62 percent in the first six months, mean time to acknowledge an alert is under five minutes, and the engineering team spends its nights sleeping instead of watching dashboards.

We needed 24/7 coverage without hiring a 24/7 team. Their SRE service gave us real monitoring, real on-call and a lot more sleep.
Chief Technology Officer, Kuda