Get in Touch

Course Outline

Core Principles of Cloud Operations on AWS

  • Defining operational roles and duties in the cloud
  • AWS account architecture, organizational setup, and multi-account approaches
  • Key operational tools: CloudWatch, CloudTrail, and AWS Config

Infrastructure as Code and Setup

  • Foundations of IaC and immutable infrastructure concepts
  • Setting up resources using Terraform and AWS CloudFormation
  • Handling state, modules, and environment transitions

CI/CD and Release Methods

  • Building CI/CD workflows for cloud-native applications
  • Blue/green, canary, and rolling release techniques
  • Automating rollbacks, health verification, and release validation

Tracking, Observability, and Notifications

  • Handling metrics, logs, and traces: collection, storage, and analysis
  • Utilizing CloudWatch, X-Ray, and external observability platforms
  • Establishing SLOs/SLIs, alerting rules, and on-call protocols

Security Management and Identity Control

  • IAM best practices, least-privilege principles, and cross-account permissions
  • Managing secrets, KMS, and secure parameter stores
  • Operational security measures: patching methods, vulnerability scanning, and audit logs

Resilience, Backups, and Disaster Recovery

  • Architecting for fault tolerance and high availability
  • Backup plans, automated snapshots, and restoration procedures
  • Disaster recovery strategies and creating operational runbooks

Cost Efficiency and Governance

  • Cost transparency: invoicing, tagging, and allocation methods
  • Right-sizing resources, reserved instances/savings plans, and budget controls
  • Governance frameworks: policies, guardrails, and compliance automation

Containers, Serverless, and Runtime Management

  • Operational needs for ECS, EKS, and Lambda
  • Service discovery, automatic scaling, and resource constraints
  • Logging, tracing, and troubleshooting containerized applications

Incident Handling, Playbooks, and Chaos Engineering

  • Runbook-based incident response and post-mortem reviews
  • Automating fixes and self-healing patterns
  • Introduction to chaos testing for verifying system resilience

Practical Session: Managing a Sample Workload

  • Deploying a test application via IaC and a CI/CD pipeline
  • Setting up monitoring, alerts, and automated fix scripts
  • Simulating failures and practicing runbook-based responses

Recap and Future Directions

Requirements

  • Foundational knowledge of cloud principles and networking
  • Competence with the Linux command line and scripting
  • Practical experience with source control (Git) and fundamental CI/CD concepts

Target Audience

  • Cloud operations specialists
  • SREs and platform engineers
  • DevOps engineers and technical team leaders
 21 Hours

Testimonials (1)

Related Categories