Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Agentic AI in Operations
- Transitioning from static runbooks to reasoning agents: The evolution of IT automation
- Anatomy of an agent: Reasoning loops, tool utilization, memory, and planning
- Determining when to automate versus when to retain human oversight
Agent Frameworks and System Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling cycles
- Multi-agent architectures: Supervisor, hierarchical, and swarm models
- Framework evaluation: LangGraph, CrewAI, AutoGen, and custom agent implementations
- Developing a basic operational agent: Querying monitoring, diagnosing, and proposing solutions
Tool Integration for IT Operations
- Linking agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-driven log querying: Integration with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: Executing kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Automating Incident Response
- Automated triage: Severity classification and ticket routing
- Generating root cause hypotheses and collecting supporting evidence
- Executing remediation: Restarting, scaling, rolling back, and failover actions
- Creating an incident runbook agent with staged autonomy levels
Safety Protocols, Guardrails, and Human Oversight
- Action categorization: Read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical tasks
- Guardrail strategies: Action allowlists, blast radius constraints, and rollback assurances
- Maintaining audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Scenarios
- Coordinating specialized agents: Triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents suggest opposing actions
- Simulating major incidents with a multi-agent response strategy
Observability and Performance Evaluation
- Tracing agent reasoning chains for debugging and auditing
- Assessing decision quality: Measuring precision, recall, and resolution time
- Feedback mechanisms: Learning from operator overrides and operational outcomes
- Tracking costs and token economics for operational agents
Production Deployment and Ongoing Operations
- Deploying agents as services: Utilizing APIs, webhooks, and scheduled tasks
- Phased autonomy rollout: Moving from shadow mode to full auto-remediation
- Managing agent failures: Procedures for when the agent itself encounters issues
- Building the business case: Measuring ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE methodologies.
- Proficiency in Python scripting and REST API interactions.
- Fundamental knowledge of LLM capabilities and prompt engineering techniques.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation strategies.
- Platform engineers developing self-healing infrastructure solutions.
- IT operations leaders assessing the potential of agentic AI for incident management.
14 Hours