Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and its significance in modern IT
- Comparing traditional monitoring with AIOps-driven observability
- Exploring AIOps architecture and essential components
Collecting and Normalizing Operational Data
- Identifying observability data types: metrics, logs, and traces
- Ingesting data from diverse sources, including servers, containers, and cloud environments
- Implementing agents and exporters such as Prometheus, Beats, and Fluentd
Data Correlation and Anomaly Detection
- Applying time series correlation and statistical methods
- Leveraging ML models for anomaly detection
- Identifying incidents across distributed systems
Alerting and Noise Reduction
- Crafting intelligent alert rules and defining effective thresholds
- Managing suppression, deduplication, and alert grouping
- Integrating with tools like Alertmanager, Slack, PagerDuty, or Opsgenie
Root Cause Analysis and Visualization
- Utilizing dashboards to visualize metrics and identify trends
- Examining events and timelines for comprehensive RCA
- Tracking issues across layers using distributed tracing tools
Automation and Remediation
- Initiating automated scripts or workflows based on incident triggers
- Connecting with ITSM systems such as ServiceNow and Jira
- Exploring use cases including self-healing, scaling, and traffic rerouting
Open Source and Commercial AIOps Platforms
- Reviewing key tools: Prometheus, Grafana, ELK, Moogsoft, and Dynatrace
- Defining evaluation criteria for selecting the right AIOps platform
- Participating in demos and hands-on exercises with a chosen stack
Summary and Next Steps
Requirements
- A solid understanding of IT operations and system monitoring concepts
- Experience working with monitoring tools or dashboards
- Familiarity with basic log and metric formats
Target Audience
- Operations teams managing infrastructure and applications
- Site Reliability Engineers (SREs)
- IT monitoring and observability teams