Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Cloud Operations on AWS
- Defining operational roles and responsibilities in the cloud.
- Understanding AWS account structures, organizations, and multi-account strategies.
- Exploring core operational services: CloudWatch, CloudTrail, and AWS Config.
Infrastructure as Code and Provisioning
- Core principles of IaC and immutable infrastructure.
- Executing provisioning tasks with Terraform and AWS CloudFormation.
- Managing state, modules, and environment promotion workflows.
CI/CD and Deployment Strategies
- Architecting CI/CD pipelines for cloud-native applications.
- Implementing blue/green, canary, and rolling deployment techniques.
- Automating rollback procedures, health checks, and release validation.
Monitoring, Observability, and Alerting
- Handling metrics, logs, and traces: shipping, storing, and analyzing data.
- Utilizing CloudWatch, X-Ray, and third-party observability platforms.
- Establishing SLOs/SLIs, alerting policies, and on-call practices.
Security Operations and Identity Management
- IAM best practices, enforcing least privilege, and managing cross-account access.
- Managing secrets, KMS, and secure parameter stores.
- Operational security measures: patching strategies, vulnerability scanning, and maintaining audit trails.
Resilience, Backup, and Disaster Recovery
- Designing systems for fault tolerance and high availability.
- Defining backup strategies, automating snapshots, and executing restore procedures.
- Planning for disaster recovery and creating effective runbooks.
Cost Optimization and Governance
- Achieving cost visibility through billing, tagging, and cost allocation strategies.
- Rightsizing resources, managing reserved instances/savings plans, and implementing budget controls.
- Establishing governance: policies, guardrails, and automation for compliance.
Containers, Serverless, and Runtime Operations
- Operational considerations for ECS, EKS, and Lambda.
- Managing service discovery, autoscaling, and resource constraints.
- Logging, tracing, and debugging containerized workloads.
Incident Response, Playbooks, and Chaos Engineering
- Executing runbook-driven incident response and postmortem analysis.
- Automating remediation and implementing self-healing patterns.
- Introduction to chaos experiments for validating system resilience.
Hands-on Workshop: Operating a Sample Workload
- Deploying a sample application using IaC and a CI/CD pipeline.
- Implementing monitoring, alerts, and automated remediation scripts.
- Simulating incidents and practicing runbook-based response procedures.
Summary and Next Steps
Requirements
- Fundamental knowledge of cloud concepts and networking.
- Proficiency with the Linux command line and scripting.
- Experience with source control systems (Git) and foundational CI/CD principles.
Target Audience
- Cloud operations engineers.
- Site Reliability Engineers (SREs) and platform engineers.
- DevOps engineers and technical team leads.
21 Hours
Testimonials (1)
I've find out new interesting things about Lambda and Serverless