This immersive, three-day workshop is dedicated to the construction and tuning of high-efficiency data processing workloads. It leverages PySpark, Pandas, and Polars within Kubernetes-based ecosystems to drive performance and reliability.
Learners will gain a robust operational understanding of Spark application execution on Kubernetes. The curriculum explores how application-level configuration choices directly impact performance metrics, scalability, resource utilization, and overall cost efficiency. Key optimization strategies are examined in depth, covering executor sizing, memory allocation, dynamic allocation, partitioning tactics, shuffle mechanics, mitigation of small-file bottlenecks, and best practices for efficient Parquet handling.
The program also tackles frequent obstacles encountered when utilizing Pandas, such as memory ceilings and out-of-memory exceptions. It presents Polars as a high-performance alternative for specific data processing tasks. Through interactive, hands-on sessions, participants will learn to diagnose performance and memory constraints, evaluate various configuration strategies, and implement optimization techniques in realistic ETL and machine learning contexts.
The core focus of the course is on practical decision-making: mastering the identification of bottlenecks, selecting the optimal tool for the job, configuring Spark for maximum efficiency, and striking the right balance between performance and infrastructure cost.
Read more...