Reliability Engineering

Fixed once, stays fixed; ready for the next failure

Makes sure a problem fixed once stays fixed.

RPR· Repeat Prevention Rate SRE on DevSemantic

For VP Engineering, Head of SRE

  • raw incident counts
  • uptime percentages on monthly reports
  • the annual disaster-recovery drill
  • threshold alerts at 90% utilisation

What changes

Stop raw incident counts. Start seeing it live.

Makes sure a problem fixed once stays fixed, and prepares services for failures that haven't happened yet. It verifies root causes and tracks whether they recur; puts every critical user journey on a service-level objective with a live error budget; tests the main failure modes of critical services and verifies recovery; and forecasts saturation of compute, storage, connections and quotas well ahead of time.

Before · reports you wait for

  • raw incident countsreport
  • uptime percentages on monthly reportsreport
  • the annual disaster-recovery drillreport
  • threshold alerts at 90% utilisationreport

After · live on DevSemantic

  • RPR Repeat Prevention RateFixed problems that stay fixed95%↑
  • EBA Error Budget AdherenceCritical services within their error budget92%↑
  • FMC Failure Mode CoverageFailure modes tested with recovery verified84%↑
DevSemantic

How it works

DevSemantic is a context plane for site reliability and operations. It connects telemetry, services, dependencies, releases and runbooks into one model, so problems are predicted from leading signals, handled by approved automation, and prevented from returning. The four solutions follow the life of a production problem: see it clearly, catch it before it pages, stop it coming back, and stop releases from introducing it.

We model how production really runs

Services, dependencies, releases, telemetry and runbooks in one live model.

What we are measured on

One headline number. Three that back it up.

RPRRepeat Prevention RateHeadline KPI

Fixed problems that stay fixed

Share of verified root causes that do not recur in the review period after their fix ships

95%

your baseline · 48 %+47 pts in 90 days

Commitment · Up against baseline; the same cause does not returnEvidence · Incident history

  1. EBAError Budget AdherenceSupporting

    Critical services within their error budget

    Share of critical services that end the review period within their error budget

    92%↑ vs baseline

    Replaces uptime percentages on a monthly reportEvidence · Customer's telemetry

  2. FMCFailure Mode CoverageSupporting

    Failure modes tested with recovery verified

    Share of critical services whose main failure modes (dependency loss, zone outage, saturation, bad config) were tested in the period, with recovery verified against target

    84%↑ vs baseline

    Replaces the annual disaster-recovery drillEvidence · Experiment records

  3. SFLSaturation Forecast LeadSupporting

    Days of warning before something saturates

    Days of warning between a capacity forecast and the resource reaching saturation, across compute, storage, connections and quotas

    21days↑ vs baseline

    Replaces threshold alerts at 90% utilisationEvidence · Incident history

Numbers shown are illustrative. Yours start from your own baseline, measured in the diagnostic.

Proof you can open

Every number comes from a record you own.

Here that record is the incident history. If we cannot show where a number came from, we do not report it.

Incident historyappend-only · yours to inspect
  1. day 1 · 09:04rcaroot cause verified · stale DNS cache on svc-auth · fix shipped r-3391
  2. day 2 · 11:21trackreview period: no recurrence · RPR sample recorded
  3. day 4 · 13:38slojourney checkout · budget 99.9% · 62% remaining
  4. day 5 · 15:55experimentzone outage test · svc-orders · recovery 04:10 vs target 05:00 · pass
  5. day 7 · 17:12forecastrds-connections saturates in 19 days · ticket opened

Questions

What people ask before they start.

Does an AI decide the root cause?

Agents assemble and correlate the evidence and propose a cause. Your pod and your engineers verify it against the evidence before it is recorded.

Who writes the fix?

Your pod, with your developers, in your repositories and review process.

How do you know a fix worked?

When the cause does not return in the review period. Until then, the fix is shipped but not verified.

What if the evidence is not there?

The analysis says so. A missing log or metric is reported as UNKNOWN, and closing the gap becomes part of the work.

See Reliability Engineering on your own estate.

A 30-minute walkthrough with an engineer. We show the RPR loop running and answer what it would look like for you.

Every solution is sold against one headline KPI, committed for 90 days against your own baseline and reported from evidence you can inspect.

  • SOC 2Type 2
  • HIPAACompliant
  • GDPRCompliant
  • ISO 270012013
  • ISO 90012015
  • ISO 200002018
  • ISO 134852016