This intensive three-day workshop is dedicated to constructing and refining high-performance data processing workloads utilizing PySpark, Pandas, and Polars within Kubernetes-based architectures.
Attendees will cultivate a deep, practical grasp of how Spark applications operate on Kubernetes, exploring how strategic configuration choices directly impact performance, scalability, resource efficiency, and overall operational costs. The curriculum delves into critical optimization domains, such as executor sizing, memory distribution, dynamic allocation, partitioning methodologies, shuffle dynamics, mitigating small-file inefficiencies, and enhancing Parquet handling.
Additionally, the course tackles prevalent challenges associated with Pandas, particularly memory constraints and out-of-memory exceptions, while introducing Polars as a robust, high-performance alternative for specific data tasks. Through immersive, hands-on labs, participants will learn to troubleshoot performance and memory bottlenecks, evaluate various configuration approaches, and implement optimization techniques in realistic ETL and machine learning contexts.
The core focus remains on practical decision-making: mastering the ability to pinpoint bottlenecks, select the right tool for the job, configure Spark effectively, and strike a balance between performance and infrastructure resource usage and cost.
Read more...