Solutions · SRE

Give your engineers their toil hours back.

Runbook Automation turns the runbooks you already have into agent playbooks that handle routine operational work under policy you set, logged and reversible, and keeps them current as your systems change.

Engineering hours returned from routine operational work.

  • More toil hours returned, counted every sprint. Measured by THR, Toil Hours Returned.
  • More of your top scenarios run as playbooks. Measured by top scenarios automated.
  • Fewer out-of-hours pages for routine work. Measured by out-of-hours toil pages.
THR · Toil Hours Returnedticket closure countsup against your baseline · reported every sprint

Supporting KPIs

Top scenarios automated ↑

Share of the highest-toil scenarios, agreed in the diagnostic, with a playbook in use

approval record

Playbook runs with a verified outcome ↑

Share of playbook runs whose result was checked against the system's state

approval record

Playbooks out of date ↓

Playbooks acting on a service that changed since they were last reviewed

context plane

Out-of-hours toil pages ↓

Pages outside working hours for scenarios a playbook covers

your records

Baselines are set in the 30-day diagnostic. Targets are committed in the 90-day scorecard. A KPI we cannot measure is reported as UNMEASURED.

A runbook that a person follows at 3am is a script waiting to be run.

Most operational work is already written down. Restart the consumer. Rotate the credential. Clear the disk. Route the ticket. An engineer reads the steps and does them, again and again, often out of hours.

Read why

The runbooks age as the systems change, so the engineer also has to know which steps are still true.

Scripting them has been tried. Scripts break when the system moves, and nobody wants a script taking actions it does not understand.

What it costs

  • Engineering hours spent on work a runbook already describes.
  • Runbooks that are wrong by the time someone follows them.
  • Routine work that interrupts the work only engineers can do.

One runbook, now a playbook.

  • runbook: settlement-worker consumer stuck · 6 steps · last followed by hand 14 times this quarter
  • playbook: the same steps, grounded in the service's dependencies and owner
  • run under policy: restart on a named service is allowed · logged · undo armed
  • outcome verified · queue draining · ticket closed with the record attached

illustrative console output · the categories and the final line are the contract

Runbooks your agents can follow, and that stay true.

Runbook Automation runs on DevSemantic, so each playbook knows the service it acts on, what depends on it and who owns it.

Existing runbooks, wikis and ticket histories are turned into agent playbooks with your team, starting with the scenarios that cost the most hours.

What you get

  • Agent playbooks for your top operational scenarios
  • A policy for what runs unattended and what waits for a person
  • A record of every run and its outcome
  • Toil hours returned, reported every sprint

Industry lens

End-of-day and settlement routines run under policy, with the record your operations risk team needs.

What it is not

Not a general scripting tool. Playbooks run only the steps your policy allows, on the systems they are grounded in.

FAQ

What if our runbooks are incomplete?

Most are. The pod works from your runbooks, ticket history and the engineers who do the work, and the playbook becomes the runbook that is kept current.

What happens if a step fails?

The playbook stops, records what happened, undoes what it can and hands the case to a person.

How are toil hours measured?

With your team, every sprint, against the hours the same work took in the diagnostic baseline.

Do we need to change our tools?

No. Playbooks act through the tools and permissions you already have.

How Runbook Automation is delivered

A capability, the pod that runs it as managed support, and a scorecard you hold us to.

  • Diagnostic · days 0 to 30, at no cost.

    DevSemantic is connected to your tickets, runbooks and on-call records. The scenarios that cost the most hours are ranked, and every KPI baseline is set.

  • Outcome · days 30 to 90.

    The top scenarios become playbooks, run first with a person approving each step, then under policy once your team approves it.

  • The Semantic Loop · from day 90.

    Every run and every system change is fed back, so playbooks stay true and new scenarios are added.

The pod

two senior engineers, run as managed support

  • Forward-deployed SRE.

    Connects DevSemantic to your systems and owns the KPI baseline with you.

  • Incident automation engineer.

    Turns runbooks into playbooks with your engineers and keeps them current.

What runs underneath

The gate holds every risky step for a person.

Count the hours your runbooks cost.

In the diagnostic we rank the operational scenarios that take the most engineering hours and show which ones a playbook could run.

  • SOC 2Type 2
  • HIPAACompliant
  • GDPRCompliant
  • ISO 270012013
  • ISO 90012015
  • ISO 200002018
  • ISO 134852016