Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open Source Tools
- Core AIOps concepts and their strategic benefits
- The role of Prometheus and Grafana within the observability stack
- Positioning ML in AIOps: contrasting predictive and reactive analytics
Setting Up Prometheus and Grafana
- Installation and configuration of Prometheus for time series data collection
- Building real-time metrics dashboards in Grafana
- Deep dive into exporters, relabeling, and service discovery mechanisms
Data Preprocessing for ML
- Extraction and transformation of Prometheus metrics
- Curating datasets optimized for anomaly detection and forecasting tasks
- Utilizing Grafana’s transformation features or Python-based pipelines
Applying Machine Learning for Anomaly Detection
- Introduction to foundational ML models for outlier detection (e.g., Isolation Forest, One-Class SVM)
- Training and evaluating models using time series datasets
- Visualizing detected anomalies directly on Grafana dashboards
Forecasting Metrics with ML
- Developing introductory forecasting models (ARIMA, Prophet, LSTM)
- Predicting trends in system load and resource consumption
- Leveraging predictions to facilitate early alerting and scaling decisions
Integrating ML with Alerting and Automation
- Crafting alert rules that respond to ML outputs or predefined thresholds
- Managing notifications and routing via Alertmanager
- Automating script execution and workflows in response to detected anomalies
Scaling and Operationalizing AIOps
- Integrating complementary observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
- Embedding ML models into broader observability pipelines for operational use
- Best practices for managing AIOps at enterprise scale
Summary and Next Steps
Requirements
- A solid grasp of system monitoring and observability fundamentals
- Practical experience working with Grafana or Prometheus
- Basic proficiency in Python and core machine learning principles
Target Audience
- Observability engineers
- Infrastructure and DevOps specialists
- Monitoring platform architects and Site Reliability Engineers (SREs)