How to evaluate an autonomous SRE
If you're evaluating an autonomous SRE for your platform team, here are the questions to ask — and the answers that should make you walk away. Some of these we pass. If a competitor passes more of them, you should buy that one.
I'm writing this for the platform leader who has been asked to evaluate Aizen alongside two or three other vendors and doesn't want to be the person who picks wrong. I'd rather lose your evaluation to a vendor who passes more of these than win it on marketing.
Eleven questions. Skip the ones that don't apply to your environment.
The right answer is your VPC, especially in a regulated industry. If telemetry leaves your environment to be processed in the vendor's cloud, that's a compliance, security, and trust issue all at once. The agent should run where the data lives.
Read-only telemetry for ingest. Scoped write access to a defined set of action endpoints — Kubernetes API, Terraform, cloud CLIs. No standing access to data stores.
Walk away if: the vendor wants broad admin access "to be safe."
Look for a defined action library — pre-approved runbooks with explicit parameters. The agent should not invent fixes. It should match the current incident to existing runbooks and apply them.
Press harder if: the vendor describes their agent as "deciding what to do." Ask exactly how the decision space is constrained.
Every action should have an explicit rollback plan, verified before execution. The agent should automatically roll back if the action doesn't restore the system within an expected window.
Walk away if: the answer is "we don't need rollback because the actions are safe."
There should be at least three: autonomous low-risk, approval-required medium-risk, and blocked-from-autonomy high-risk. The boundaries should be configurable per team and per runbook.
Walk away if: there's a single autonomous yes/no toggle. That vendor hasn't met operational reality.
Every action, every decision, every verification. Inputs, outputs, parameters, outcomes. The log should be queryable and tamper-resistant. Ask specifically about log retention and chain-of-custody guarantees.
The agent should not attempt to autonomously resolve incidents that don't match its runbook library. It should join the human-led response as a contributor: pulling logs, surfacing precedents, drafting the post-mortem.
Walk away if: the vendor claims their agent resolves anything autonomously. That's fiction.
After a novel incident is resolved, the agent should propose a new runbook for team review. Your team approves or rejects. The library grows with your explicit consent. Automatic learning without human review is a no.
Read access to your observability stack — Splunk, Prometheus, CloudWatch, or whatever you run. Scoped write access to your control plane. A chat integration. No new instrumentation, no agents on hosts.
Scrutinise if: you need to deploy new infrastructure just to use the product.
Read-only ingest should take a day or two. First low-risk autonomous resolutions within two to four weeks. Full autonomy across your recurring incident catalogue in three to six months.
Doubt both ends: "value in week one" is overselling. Six months to first value is underdeveloped.
Any vendor with production customers should be able to do this in under five minutes. If they can't, you're a pilot customer rather than their fifth. That's fine — but you should know it, and price the deal accordingly.
How Aizen scores on its own checklist
Honestly:
On question 11: today we'd show you sample logs from our own dogfooding environment, not a customer incident. If you become a design partner, you'll have a real audit log to show your next vendor evaluation.
One more thing worth saying plainly, because it will come up in your security review: we're working toward SOC 2 Type II and are not certified today. Some vendors in this category are. If that's a hard requirement for your organisation right now, that's a reasonable place to rule us out, and I'd rather you learn it here than three calls in.
If a competitor passes 11 of 11 and we pass 10, buy theirs. The point of this checklist is to make sure you buy a system you can trust in production — not to make sure you buy ours.
But ask them. Ask all of them. The vendors who can't answer these cleanly are the ones who'll cost you a 4 AM outage in eighteen months. The vendors who can are the ones worth piloting.