After assistants that answer questions, it is now the turn of agents: AI systems that observe, decide and take action in real tools. In IT operations, that means an agent that reads alerts, checks logs, runs a runbook and updates the incident without waiting for someone to open the console. The promise is big. So is the risk. Anyone who runs Command Centers for large banks and payment providers knows that one wrong action at three in the morning costs more than ten unanswered alerts.
Where you can already trust them
Agents work well on tasks that are low risk, highly repetitive and easy to verify:
- Alert correlation: grouping events from the same incident, removing duplicates and pointing to the likely source cuts the noise that reaches the NOC.
- Initial diagnosis: collecting metrics, logs and recent changes and handing the analyst a ready summary instead of a blank screen.
- Reversible runbooks: restarting a service, clearing a temp directory, adding a replica. Known, documented actions that can be undone.
- Logging and communication: opening the incident, filling in the timeline and drafting the status update for review.
Where a human is still needed
Anything irreversible, ambiguous or high impact still needs a person making the call: failover between data centers, changes to production databases, access changes, rollbacks that involve data migrations and any external communication to customers or regulators. Novel incidents belong here too. Agents are good at recognizing patterns they have seen before; faced with the unknown, they tend to force it into what they already know.
Autonomy is earned one category of action at a time, not through enthusiasm for the technology.
Guardrails you can’t skip
- A catalog of allowed actions: the agent only executes what is explicitly listed, with defined scope, environment and limits. Everything else is denied by default.
- Approval tiers: low-risk actions run on their own; medium-risk ones ask for one-click confirmation; critical ones require approval from a named owner.
- Least-privilege credentials: the agent has its own identity with narrow, time-bound permissions, never an administrator’s account.
- A complete audit trail: what the agent saw, what it decided, why, and what it executed. Without it there is no root cause analysis and no compliance.
- A kill switch: any operator must be able to suspend the agent immediately, without depending on whoever configured it.
How to get started
Start in suggestion mode: the agent proposes and the analyst executes. Track the hit rate by type of action for a few weeks. Promote to automatic execution only the categories that have proven reliable, and review that list regularly, like any other change to operations.
Well governed, AI agents don’t replace the NOC. They take the mechanical work off the team and give back time for what requires judgment: understanding the problem, deciding and communicating.
Want to talk about this?
Tell us about your challenge. The conversation is direct and with no commitment.