Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and understanding its significance
- Comparing traditional monitoring with AIOps-driven observability
- Examining AIOps architecture and core components
Collecting and Normalizing Operational Data
- Understanding types of observability data: metrics, logs, and traces
- Ingesting data from diverse sources such as servers, containers, and the cloud
- Utilizing agents and exporters (Prometheus, Beats, Fluentd)
Data Correlation and Anomaly Detection
- Applying time series correlation and statistical methods
- Leveraging ML models for effective anomaly detection
- Identifying incidents across distributed systems
Alerting and Noise Reduction
- Developing intelligent alert rules and thresholds
- Implementing suppression, deduplication, and alert grouping techniques
- Integrating with platforms like Alertmanager, Slack, PagerDuty, or Opsgenie
Root Cause Analysis and Visualization
- Visualizing metrics and detecting trends using dashboards
- Analyzing events and timelines to facilitate RCA
- Tracing issues across various layers using distributed tracing tools
Automation and Remediation
- Initiating automated scripts or workflows in response to incidents
- Integrating with ITSM systems such as ServiceNow or Jira
- Exploring use cases including self-healing, scaling, and traffic rerouting
Open Source and Commercial AIOps Platforms
- Surveying available tools: Prometheus, Grafana, ELK, Moogsoft, Dynatrace
- Establishing evaluation criteria for selecting an AIOps platform
- Conducting demonstrations and hands-on activities with a selected stack
Summary and Next Steps
Requirements
- A solid grasp of IT operations and system monitoring concepts
- Practical experience with monitoring tools or dashboards
- Familiarity with standard log and metric formats
Target Audience
- Operations teams managing infrastructure and application environments
- Site Reliability Engineers (SREs)
- Teams dedicated to IT monitoring and observability