Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop ecosystem
  • Brief overview of Python and Scala

Core Concepts (Theory):

  • System architecture
  • Resilient Distributed Datasets (RDDs)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Exploring fundamentals in a Databricks environment (Hands-on Workshop):

  • Practical exercises with the RDD API
  • Core action and transformation functions
  • Working with PairRDDs
  • Implementing Joins
  • Caching strategies
  • Practical exercises with the DataFrame API
  • Using SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • User Defined Functions (UDFs)
  • Introduction to the DataSet API
  • Structured Streaming

Mastering deployment in an AWS environment (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Comparing AWS EMR and AWS Glue
  • Executing sample jobs in both environments
  • Evaluating advantages and limitations

Additional Topics:

  • Introduction to Apache Airflow orchestration

Requirements

Programming proficiency (ideally in Python or Scala)

Foundational knowledge of SQL

 21 Hours

Testimonials (3)

Related Categories