Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- The evolution from static runbooks to intelligent reasoning agents in IT automation
- Anatomy of an agent: reasoning loops, tool utilization, memory management, and planning
- Determining when to automate versus when to retain human intervention
Agent Frameworks and System Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and iterative tool-calling loops
- Multi-agent structures: supervisor, hierarchical, and swarm coordination patterns
- Comparative analysis of frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Constructing your first operational agent: querying monitors, diagnosing issues, and proposing solutions
Integrating Tools for IT Operations
- Connecting agents to APIs such as Prometheus, Grafana, Datadog, and PagerDuty
- Agent-driven log querying using Elasticsearch, Loki, and Splunk integrations
- Utilizing infrastructure tools like kubectl, Terraform, and Ansible through agent actions
- Designing secure tool interfaces with parameter validation and idempotent operations
Automating Incident Response
- Automated triage processes: severity classification and intelligent routing
- Generating root cause hypotheses and gathering supporting evidence
- Executing automated remediation tasks: restarts, scaling, rollbacks, and failovers
- Developing an incident runbook agent with adjustable autonomy levels
Safety Protocols, Guardrails, and Human Oversight
- Classifying actions by risk level: read-only, low-risk, high-risk, and destructive
- Implementing approval gates and escalation policies for critical operations
- Applying guardrail strategies: action allowlists, blast radius constraints, and rollback assurances
- Maintaining audit trails and decision provenance for regulatory compliance
Orchestrating Multi-Agent Systems for Complex Incidents
- Coordinating specialized agents: triage, diagnosis, and remediation specialists
- Managing inter-agent communication and shared context
- Resolving conflicts when agents suggest contradictory actions
- Simulating end-to-end major incidents with multi-agent coordination
Observability and Performance Evaluation
- Tracing agent reasoning chains to facilitate debugging and auditing
- Assessing decision quality through metrics like precision, recall, and resolution time
- Implementing feedback loops to learn from operator overrides and outcomes
- Tracking costs and analyzing token economics for operational agents
Production Deployment and Ongoing Operations
- Deploying agents as services using APIs, webhooks, and scheduled jobs
- Implementing gradual autonomy rollouts from shadow mode to full auto-remediation
- Preparing runbooks for agent failures: handling scenarios where the agent itself malfunctions
- Building the business case and measuring ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE methodologies.
- Proficiency in Python scripting and working with REST APIs.
- A foundational understanding of Large Language Model (LLM) capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers looking to leverage AI-driven automation.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders evaluating agentic AI for improved incident management.
14 Hours