Buyer's checklist

How to evaluate an autonomous SRE

Chandni Singh · Founder, Aizen · 9 min read

If you're evaluating an autonomous SRE for your platform team, here are the questions to ask — and the answers that should make you walk away. Some of these we pass. If a competitor passes more of them, you should buy that one.

I'm writing this for the platform leader who has been asked to evaluate Aizen alongside two or three other vendors and doesn't want to be the person who picks wrong. I'd rather lose your evaluation to a vendor who passes more of these than win it on marketing.

Eleven questions. Skip the ones that don't apply to your environment.

01
Where does the agent run — in our VPC, or in the vendor's cloud?

The right answer is your VPC, especially in a regulated industry. If telemetry leaves your environment to be processed in the vendor's cloud, that's a compliance, security, and trust issue all at once. The agent should run where the data lives.

02
What access does the agent need to my systems?

Read-only telemetry for ingest. Scoped write access to a defined set of action endpoints — Kubernetes API, Terraform, cloud CLIs. No standing access to data stores.

Walk away if: the vendor wants broad admin access "to be safe."

03
How are autonomous actions bounded?

Look for a defined action library — pre-approved runbooks with explicit parameters. The agent should not invent fixes. It should match the current incident to existing runbooks and apply them.

Press harder if: the vendor describes their agent as "deciding what to do." Ask exactly how the decision space is constrained.

04
What's the rollback story?

Every action should have an explicit rollback plan, verified before execution. The agent should automatically roll back if the action doesn't restore the system within an expected window.

Walk away if: the answer is "we don't need rollback because the actions are safe."

05
How are risk tiers handled?

There should be at least three: autonomous low-risk, approval-required medium-risk, and blocked-from-autonomy high-risk. The boundaries should be configurable per team and per runbook.

Walk away if: there's a single autonomous yes/no toggle. That vendor hasn't met operational reality.

06
What does the audit log capture?

Every action, every decision, every verification. Inputs, outputs, parameters, outcomes. The log should be queryable and tamper-resistant. Ask specifically about log retention and chain-of-custody guarantees.

07
How do you handle novel incidents?

The agent should not attempt to autonomously resolve incidents that don't match its runbook library. It should join the human-led response as a contributor: pulling logs, surfacing precedents, drafting the post-mortem.

Walk away if: the vendor claims their agent resolves anything autonomously. That's fiction.

08
How does the agent learn from new incidents?

After a novel incident is resolved, the agent should propose a new runbook for team review. Your team approves or rejects. The library grows with your explicit consent. Automatic learning without human review is a no.

09
What integrations does it require?

Read access to your observability stack — Splunk, Prometheus, CloudWatch, or whatever you run. Scoped write access to your control plane. A chat integration. No new instrumentation, no agents on hosts.

Scrutinise if: you need to deploy new infrastructure just to use the product.

10
What's the deployment timeline?

Read-only ingest should take a day or two. First low-risk autonomous resolutions within two to four weeks. Full autonomy across your recurring incident catalogue in three to six months.

Doubt both ends: "value in week one" is overselling. Six months to first value is underdeveloped.

11
Show me an audit log from a real customer incident, with PII redacted.

Any vendor with production customers should be able to do this in under five minutes. If they can't, you're a pilot customer rather than their fifth. That's fine — but you should know it, and price the deal accordingly.

How Aizen scores on its own checklist

Honestly:

Questions 1–10 · we pass, by design
Question 11 · not yet — we're onboarding three design partners this quarter

On question 11: today we'd show you sample logs from our own dogfooding environment, not a customer incident. If you become a design partner, you'll have a real audit log to show your next vendor evaluation.

One more thing worth saying plainly, because it will come up in your security review: we're working toward SOC 2 Type II and are not certified today. Some vendors in this category are. If that's a hard requirement for your organisation right now, that's a reasonable place to rule us out, and I'd rather you learn it here than three calls in.

If a competitor passes 11 of 11 and we pass 10, buy theirs. The point of this checklist is to make sure you buy a system you can trust in production — not to make sure you buy ours.

But ask them. Ask all of them. The vendors who can't answer these cleanly are the ones who'll cost you a 4 AM outage in eighteen months. The vendors who can are the ones worth piloting.

Use this checklist to evaluate us. If we don't pass, tell us — hello@aizenops.ai