Get in Touch
 Duration 21 hours

Course Outline

Introduction:

  • Apache Spark within the Hadoop ecosystem
  • Foundational overview of Python and Scala

Core Concepts (Theory):

  • Spark architecture
  • RDDs (Resilient Distributed Datasets)
  • Transformations and actions
  • Stages, tasks, and dependencies

Applying Fundamentals in Databricks (Hands-on Workshop):

  • Practical exercises with the RDD API
  • Core action and transformation functions
  • Working with PairRDDs
  • Join operations
  • Caching strategies
  • Practical exercises with the DataFrame API
  • SparkSQL integration
  • DataFrame operations: select, filter, group, and sort
  • Implementing UDFs (User Defined Functions)
  • Exploring the DataSet API
  • Streaming data

Deployment Strategies in AWS (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Comparing AWS EMR and AWS Glue
  • Building example jobs in both environments
  • Evaluating advantages and disadvantages

Additional Topics:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming proficiency (ideally in Python or Scala)

Familiarity with SQL fundamentals

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories