Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Core Principles of Cloud Operations on AWS
- Defining operational roles and duties in the cloud
- AWS account architecture, organizational setup, and multi-account approaches
- Key operational tools: CloudWatch, CloudTrail, and AWS Config
Infrastructure as Code and Setup
- Foundations of IaC and immutable infrastructure concepts
- Setting up resources using Terraform and AWS CloudFormation
- Handling state, modules, and environment transitions
CI/CD and Release Methods
- Building CI/CD workflows for cloud-native applications
- Blue/green, canary, and rolling release techniques
- Automating rollbacks, health verification, and release validation
Tracking, Observability, and Notifications
- Handling metrics, logs, and traces: collection, storage, and analysis
- Utilizing CloudWatch, X-Ray, and external observability platforms
- Establishing SLOs/SLIs, alerting rules, and on-call protocols
Security Management and Identity Control
- IAM best practices, least-privilege principles, and cross-account permissions
- Managing secrets, KMS, and secure parameter stores
- Operational security measures: patching methods, vulnerability scanning, and audit logs
Resilience, Backups, and Disaster Recovery
- Architecting for fault tolerance and high availability
- Backup plans, automated snapshots, and restoration procedures
- Disaster recovery strategies and creating operational runbooks
Cost Efficiency and Governance
- Cost transparency: invoicing, tagging, and allocation methods
- Right-sizing resources, reserved instances/savings plans, and budget controls
- Governance frameworks: policies, guardrails, and compliance automation
Containers, Serverless, and Runtime Management
- Operational needs for ECS, EKS, and Lambda
- Service discovery, automatic scaling, and resource constraints
- Logging, tracing, and troubleshooting containerized applications
Incident Handling, Playbooks, and Chaos Engineering
- Runbook-based incident response and post-mortem reviews
- Automating fixes and self-healing patterns
- Introduction to chaos testing for verifying system resilience
Practical Session: Managing a Sample Workload
- Deploying a test application via IaC and a CI/CD pipeline
- Setting up monitoring, alerts, and automated fix scripts
- Simulating failures and practicing runbook-based responses
Recap and Future Directions
Requirements
- Foundational knowledge of cloud principles and networking
- Competence with the Linux command line and scripting
- Practical experience with source control (Git) and fundamental CI/CD concepts
Target Audience
- Cloud operations specialists
- SREs and platform engineers
- DevOps engineers and technical team leaders
21 Hours
Testimonials (1)
I've find out new interesting things about Lambda and Serverless