Course Outline
Introduction
Concepts of Big Data
Introduction to Spark
Python Fundamentals
PySpark Overview
- Managing Data Distribution via the Resilient Distributed Datasets Framework
- Distributing Computation Tasks Using Spark API Operators
Integrating Python with Spark
Configuring the PySpark Environment
Deploying Spark on Amazon Web Services (AWS) EC2 Instances
Setting Up the Databricks Platform
Configuring an AWS EMR Cluster
Foundations of Python Programming
- Initiating Python Development
- Utilizing the Jupyter Notebook Interface
- Managing Variables and Primitive Data Types
- Handling List Structures
- Implementing Conditional Logic with if Statements
- Processing User Input
- Controlling Flow with while Loops
- Defining and Using Functions
- Object-Oriented Programming with Classes
- File Handling and Exception Management
- Working with Projects, Datasets, and APIs
Essentials of Spark DataFrames
- Initiating Work with Spark DataFrames
- Executing Core Operations in Spark
- Performing Grouping and Aggregation Tasks
- Processing Timestamps and Date Values
Practical Spark DataFrame Project
Machine Learning Concepts with MLlib
Applying MLlib, Spark, and Python to Machine Learning Tasks
Regressive Analysis
- Foundations of Linear Regression Theory
- Writing Code for Regression Evaluation
- Executing a Linear Regression Practical Task
- Understanding Logistic Regression Theory
- Developing Logistic Regression Algorithms
- Executing a Logistic Regression Practical Task
Random Forests and Decision Trees
- Theory Behind Tree-Based Methods
- Coding Decision Trees and Random Forest Models
- Executing a Random Forest Classification Practical Task
K-means Clustering Application
- Theoretical Background of K-means Clustering
- Developing K-means Clustering Algorithms
- Executing a Clustering Practical Task
Developing Recommender Systems
Implementing Natural Language Processing
- Principles of Natural Language Processing (NLP)
- Survey of NLP Toolsets
- Executing a Natural Language Processing Practical Task
Real-Time Streaming with Spark and Python
- Introduction to Spark Streaming Capabilities
- Executing a Spark Streaming Practical Task
Requirements
- Fundamental programming proficiency
Intended Audience
- Software Developers
- IT Specialists
- Data Scientists
Testimonials (6)
I liked that it was practical. Loved to apply the theoretical knowledge with practical examples.
Aurelia-Adriana - Allianz Services Romania
Course - Python and Spark for Big Data (PySpark)
The course was about a series of very complex related topics & Pablo has in-depth expertise of each of them. Sometimes nuances were lost in communication and/or due to time pressures and possibly expectations were not quite met due to this. Also there were some UHG/Azure Databricks setup issues however Pablo / UHG resolved these quickly once they became apparent - this to me showed a high level of understanding and professionalism between UHG & Pablo,
Michael Monks - Tech NorthWest Skillnet
Course - Python and Spark for Big Data (PySpark)
Individual attention.
ARCHANA ANILKUMAR - PPL
Course - Python and Spark for Big Data (PySpark)
Hands on Training..
Abraham Thomas - PPL
Course - Python and Spark for Big Data (PySpark)
The lessons were taught in a Jupyter notebook. The topics were structured with a logical sequence and naturally helped develop the session from the easier parts to the more complex. I'm already an advanced user of Python with background in Machine Learning, so found the course easier to follow than, possibly, some of my classmates that took the training course. I appreciate that some of the most elementary concepts were skipped and that he focused on the most substantial matters.
Angela DeLaMora - ADT, LLC
Course - Python and Spark for Big Data (PySpark)
practice tasks