Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Introduction:
- Apache Spark within the Hadoop ecosystem
- Foundational overview of Python and Scala
Core Concepts (Theory):
- Spark architecture
- RDDs (Resilient Distributed Datasets)
- Transformations and actions
- Stages, tasks, and dependencies
Applying Fundamentals in Databricks (Hands-on Workshop):
- Practical exercises with the RDD API
- Core action and transformation functions
- Working with PairRDDs
- Join operations
- Caching strategies
- Practical exercises with the DataFrame API
- SparkSQL integration
- DataFrame operations: select, filter, group, and sort
- Implementing UDFs (User Defined Functions)
- Exploring the DataSet API
- Streaming data
Deployment Strategies in AWS (Hands-on Workshop):
- Fundamentals of AWS Glue
- Comparing AWS EMR and AWS Glue
- Building example jobs in both environments
- Evaluating advantages and disadvantages
Additional Topics:
- Introduction to Apache Airflow for orchestration
Requirements
Programming proficiency (ideally in Python or Scala)
Familiarity with SQL fundamentals
Testimonials (3)
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
1. Right balance between high level concepts and technical details. 2. Andras is very knowledgeable about his teaching. 3. Exercise
Steven Wu - Intelligent Medical Objects
Course - Apache Spark in the Cloud
Get to learn spark streaming , databricks and aws redshift