Skip to content

AI SRE · Incident response · Evaluation

What AI SRE tools do, and how to evaluate one

AI SRE tools triage alerts, correlate deploys with errors and propose fixes. What they do, where they stop, and a two-week plan for evaluating one.

By
Escanor team
Published
Reading time
6 min read
On this page

An AI SRE tool is software that uses a language model to do part of the operations work a site reliability engineer does: triaging alerts, investigating incidents, connecting a failure to the deploy or config change that caused it, drafting a fix or rollback, and writing up what happened. The tool reads from systems you already run (code host, deploy platform, logs, error tracker, pager) and, in some products, acts on them under approval. Products differ less in what they claim to do than in how they show evidence, how much they are allowed to change, and how easily you can stop them.

This guide covers what these tools do in practice, where they stop being useful, and how to test one against your own incidents before you trust it.

What AI SRE tools do

Most products in this category combine four jobs. Some do all four; many start with the first two.

Triage and correlation

An alert on its own says something is wrong. A useful first answer says what changed just before it went wrong. The tool pulls recent deploys, merged pull requests, config changes and error spikes into one timeline for the affected service, so the on-call engineer starts from "the 14:02 deploy of the checkout service" instead of from a blank dashboard.

The quality of this step depends on which sources are connected. A tool that sees your error tracker but not your deploy history cannot connect the two.

Investigation

Given an incident, the tool reads logs, traces and recent changes and proposes a probable cause. Good tools cite the specific log lines, commits or deploys behind that guess. Weak ones produce a fluent paragraph with no references, and you cannot tell a correct diagnosis from a confident wrong one.

Treat any generated root cause as a hypothesis. The useful output is the evidence trail it hands you, because that is what you can check.

Remediation

Some tools go further and propose an action: roll back a deploy, revert a commit, restart a service, scale a resource, or open a pull request with a code fix. Products split on how that action runs:

ModelWhat happensWhat to check
Suggest onlyThe tool writes the command or patch; a person runs it.Whether suggestions are specific enough to run.
Approve then actThe tool prepares the action and waits for a named person to approve it.Who can approve, from where, and what the approver sees.
Act within policyThe tool runs pre-approved action types on its own and asks for the rest.How the policy is defined, and what happens with an action it does not recognise.

None of these is right for every team. A team with a well-tested rollback path may be comfortable letting a tool roll back automatically. The same team may want every database change approved by hand.

Toil and write-ups

Google's SRE book defines toil as work "tied to running a production service that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows" (Eliminating Toil). Gathering the same five dashboards for every page fits that definition, and it is where these tools save the most time.

The same book describes a postmortem as "a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions" (Postmortem Culture). A tool that already holds the incident timeline can draft that record. Someone who was there still has to correct it.

Where they stop

These limits apply to every tool in the category, including ours.

  • They only see what is connected. If your logs live somewhere the tool cannot read, its analysis has a gap it may not tell you about. A quiet dashboard is not proof of a healthy service.
  • A generated cause is unconfirmed. The model can link two events that happened close together without one causing the other.
  • The provider stays the source of truth. Your deploy platform, cloud console and pager record what happened. Check them before you close an incident.
  • Agency is a risk of its own. OWASP lists "Excessive Agency" as a top risk for LLM applications and traces it to three causes: excessive functionality, excessive permissions and excessive autonomy (OWASP LLM06:2025). An AI SRE tool with write access to production is exactly the kind of system that guidance is about.

How to evaluate an AI SRE tool

Score each candidate against the same criteria. The questions in the right-hand column are the ones a demo tends to skip.

CriterionWhat good looks likeAsk or test
Data sourcesConnects to the code host, deploy platform, logs, error tracker and pager you use today.Which of our sources are read-only, which are writable, which are missing?
EvidenceEvery claim links to a log line, deploy, commit or alert.Pick a claim at random and follow the link. Does it support the claim?
Action modelReads, writes and destructive actions are separate tiers with separate rules.What happens when it tries an action type nobody configured?
ApprovalsA named person approves risky actions and sees the exact target before approving.Can it approve its own action? Can an approver act from a phone?
CredentialsThe model never sees raw provider secrets.Where are tokens stored, and are they ever placed in a prompt or tool argument?
AuditEvery action is logged with who asked, what ran and the result.Reconstruct yesterday's actions from the log alone.
Stop controlOne switch pauses all automated work.Press it during a test run. What was already submitted?
Agent accessYour coding agents can reach the same tools under the same rules.Does it expose an MCP server or API, and with what scopes?
CostPricing you can predict from your incident volume and team size.What counts as billable usage?

A two-week evaluation plan

A demo shows the tool on the vendor's incidents. This plan tests it on yours.

  1. Pick three resolved incidents with written postmortems and known causes. Include one where the first guess during the incident was wrong.
  2. Connect sources read-only. Use the narrowest provider permissions that let the tool read deploys, logs and alerts. Do not grant write access yet.
  3. Ask it to investigate each incident as if it were new, then compare its answer with the postmortem. Record whether it found the cause, whether it cited evidence, and whether it said so when evidence was thin.
  4. Follow five citations by hand. If more than one does not support the claim it is attached to, stop the evaluation here.
  5. Test a refusal. Ask for a destructive action on a resource it should not touch. A good tool refuses or waits for approval and records the refusal.
  6. Test one bounded write on staging with approval required: a restart, a rollback to a known-good deploy, or a pull request. Check the approver saw the exact target.
  7. Read the audit log and confirm you can reconstruct steps 3 to 6 from it.
  8. Revoke access and confirm the tool can no longer act.

If a product passes steps 3 to 8 on your own data, its marketing claims matter much less.

How Escanor approaches it

Escanor is an operational layer over the tools you already use. Here is what it does for incident work, stated against the criteria above. Details are in the linked docs.

  • Incidents gathers findings from connected providers. A summary is generated from the incident's evidence; when model-backed summarisation is unavailable it falls back to an evidence-based summary, and the docs label any probable cause as unconfirmed. See Incidents and investigation.
  • Repairs (Auto-fix, repair missions and playbooks) let you inspect the proposed target, branch, commands, tests and approval state. Availability depends on connected providers, repositories, the configured assistant and your plan, and opening a repair screen does not mean every incident can be repaired.
  • Automations run goal-driven work with configurable autonomy levels, per-action overrides, budgets and an approval queue. Stop everything now pauses automation; it does not reverse work already submitted to a provider. See Automations and Autopilot.
  • Agent access goes through one MCP server with revocable keys and separate read and invoke scopes. See Escanor MCP.
  • Limits are written down. Escanor can only collect what connected sources expose, and some incident types carry only the provider's headline.

For the incident workflow end to end, read AI incident response. Plans and limits are on the pricing page. Whichever tool you evaluate, run the eight steps above on three of your own incidents before you give it write access.