This intensive three-day workshop is designed to help you build and refine high-performance data-processing workloads by leveraging PySpark, Pandas and Polars within Kubernetes-based environments.
Attendees will gain a functional grasp of how Spark applications operate on Kubernetes, and how specific configuration choices at the application level directly impact performance, scalability, resource utilisation and overall costs. Key optimisation topics covered include executor sizing, memory allocation strategies, dynamic allocation, partitioning logic, shuffle mechanics, resolving small-file issues and enhancing Parquet processing efficiency.
Additionally, the course tackles frequent challenges encountered with Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a high-performance solution for specific data-processing tasks. Through practical, hands-on labs, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration approaches and implement optimisation techniques in realistic ETL and machine learning contexts.
Throughout the training, the focus remains on practical decision-making: learning how to pinpoint bottlenecks, choose the right tools, configure Spark for maximum efficiency and strike a balance between performance and infrastructure costs.
Read more...