Get in Touch

Course Outline

Introduction to Agentic AI in Operations

  • The evolution from static runbooks to intelligent reasoning agents in IT automation
  • Anatomy of an agent: reasoning loops, tool utilization, memory management, and planning
  • Determining when to automate versus when to retain human intervention

Agent Frameworks and System Architectures

  • Single-agent patterns: ReAct, Plan-and-Execute, and iterative tool-calling loops
  • Multi-agent structures: supervisor, hierarchical, and swarm coordination patterns
  • Comparative analysis of frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
  • Constructing your first operational agent: querying monitors, diagnosing issues, and proposing solutions

Integrating Tools for IT Operations

  • Connecting agents to APIs such as Prometheus, Grafana, Datadog, and PagerDuty
  • Agent-driven log querying using Elasticsearch, Loki, and Splunk integrations
  • Utilizing infrastructure tools like kubectl, Terraform, and Ansible through agent actions
  • Designing secure tool interfaces with parameter validation and idempotent operations

Automating Incident Response

  • Automated triage processes: severity classification and intelligent routing
  • Generating root cause hypotheses and gathering supporting evidence
  • Executing automated remediation tasks: restarts, scaling, rollbacks, and failovers
  • Developing an incident runbook agent with adjustable autonomy levels

Safety Protocols, Guardrails, and Human Oversight

  • Classifying actions by risk level: read-only, low-risk, high-risk, and destructive
  • Implementing approval gates and escalation policies for critical operations
  • Applying guardrail strategies: action allowlists, blast radius constraints, and rollback assurances
  • Maintaining audit trails and decision provenance for regulatory compliance

Orchestrating Multi-Agent Systems for Complex Incidents

  • Coordinating specialized agents: triage, diagnosis, and remediation specialists
  • Managing inter-agent communication and shared context
  • Resolving conflicts when agents suggest contradictory actions
  • Simulating end-to-end major incidents with multi-agent coordination

Observability and Performance Evaluation

  • Tracing agent reasoning chains to facilitate debugging and auditing
  • Assessing decision quality through metrics like precision, recall, and resolution time
  • Implementing feedback loops to learn from operator overrides and outcomes
  • Tracking costs and analyzing token economics for operational agents

Production Deployment and Ongoing Operations

  • Deploying agents as services using APIs, webhooks, and scheduled jobs
  • Implementing gradual autonomy rollouts from shadow mode to full auto-remediation
  • Preparing runbooks for agent failures: handling scenarios where the agent itself malfunctions
  • Building the business case and measuring ROI for autonomous operations

Requirements

  • Practical experience in IT operations, DevOps, or SRE methodologies.
  • Proficiency in Python scripting and working with REST APIs.
  • A foundational understanding of Large Language Model (LLM) capabilities and prompt engineering.

Target Audience

  • SRE and DevOps engineers looking to leverage AI-driven automation.
  • Platform engineers focused on building self-healing infrastructure.
  • IT operations leaders evaluating agentic AI for improved incident management.
 14 Hours

Related Categories