Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps Using Open Source Solutions
- Key concepts and advantages of AIOps
- The role of Prometheus and Grafana within the observability architecture
- The position of ML in AIOps: contrasting predictive and reactive analysis
Configuring Prometheus and Grafana
- Deployment and setup of Prometheus for time-series data capture
- Designing Grafana dashboards powered by live metrics
- Investigation of exporters, relabeling, and service discovery mechanisms
Data Preparation for Machine Learning
- Extraction and transformation of Prometheus metrics
- Structuring datasets for anomaly detection and forecasting tasks
- Leveraging Grafana’s built-in transformations or Python-based pipelines
Utilizing Machine Learning for Anomaly Detection
- Introduction to foundational ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
- Model training and assessment on time-series datasets
- Representing anomalies through Grafana dashboard visualizations
Metric Forecasting via Machine Learning
- Construction of basic forecasting models (ARIMA, Prophet, introductory LSTM)
- Projection of system load or resource consumption trends
- Leveraging predictions for proactive alerting and scaling strategies
Connecting ML to Alerting and Automation
- Establishing alert rules derived from ML outputs or predefined thresholds
- Implementation of Alertmanager and notification routing logic
- Initiation of scripts or automation workflows upon anomaly detection
Scaling and Operationalizing AIOps
- Integration with external observability platforms (e.g., ELK stack, Moogsoft, Dynatrace)
- Embedding ML models into observability workflows
- Recommended best practices for large-scale AIOps deployment
Conclusion and Recommended Next Steps
Requirements
- A solid grasp of system monitoring and observability fundamentals
- Prior experience with either Grafana or Prometheus
- Proficiency in Python along with an understanding of basic machine learning concepts
Target Audience
- Observability engineers
- Infrastructure and DevOps personnel
- Monitoring platform architects and Site Reliability Engineers (SREs)