Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Objectives, and Migration Strategy
- Course goals, alignment with participant profiles, and success criteria
- High-level migration approaches and associated risk factors
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Core Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Differences between SMP and MPP and their impact on migration
- Medallion (Bronze→Silver→Gold) architecture design and Unity Catalog introduction
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure into a notebook
- Mapping temporary tables and cursors to DataFrame transformations
- Validation and comparison against original outputs
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning techniques
Day 2 Lab — Incremental Ingestion & Optimization
- Implementation of Auto Loader ingestion and MERGE workflows
- Application of OPTIMIZE, Z-ORDER, and VACUUM with result validation
- Measurement of read/write performance enhancements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL capabilities: window functions, higher-order functions, and JSON/array processing
- Interpreting Spark UI, DAGs, shuffles, stages, and tasks for bottleneck diagnosis
- Query tuning strategies: broadcast joins, hints, caching, and reducing spill
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Utilizing Spark UI traces to identify and resolve skew and shuffle issues
- Benchmarking before/after performance and documenting tuning steps
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: drivers, executors, lazy evaluation, and partitioning
- Converting loops and cursors into vectorized DataFrame operations
- Modularization, UDFs/pandas UDFs, widgets, and creating reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Refactoring procedural ETL scripts into modular PySpark notebooks
- Incorporating parametrization, unit-style tests, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error handling
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment notes
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, lineage, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook development
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and key takeaways
- Gap analysis, recommended follow-up activities, and handover of training materials
- References, further learning pathways, and support options
Requirements
- Familiarity with core data engineering concepts
- Practical experience with SQL and stored procedures (e.g., Synapse or SQL Server)
- Understanding of ETL orchestration principles (e.g., ADF or equivalent tools)
Target Audience
- Technology managers with a data engineering background
- Data engineers transitioning from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing the adoption of Databricks