Course Outline
Introduction, Objectives, and Migration Strategy
- Defining course goals, aligning with participant profiles, and establishing success criteria
- Overview of high-level migration approaches and associated risk factors
- Configuring workspaces, repositories, and laboratory datasets
Day 1 — Migration Fundamentals and Architecture
- Core Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Differences between SMP and MPP and their impact on migration
- Medallion (Bronze→Silver→Gold) design principles and Unity Catalog introduction
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure into a notebook
- Mapping temporary tables and cursors to DataFrame transformations
- Validating results and comparing them against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel capabilities
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Optimization techniques: OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning
Day 2 Lab — Incremental Ingestion & Optimization
- Implementing Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM commands; verifying outcomes
- Evaluating read/write performance enhancements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL features: window functions, higher-order functions, and JSON/array processing
- Interpreting Spark UI, DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
- Query optimization patterns: broadcast joins, hints, caching, and minimizing spills
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a resource-intensive SQL process into optimized Spark SQL
- Leveraging Spark UI traces to pinpoint and resolve skew and shuffle issues
- Benchmarking before/after metrics and documenting tuning procedures
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Modularization, UDFs/pandas UDFs, widgets, and building reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Rebuilding a procedural ETL script into modular PySpark notebooks
- Incorporating parametrization, unit-style testing, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines incorporating quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retry mechanisms, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, lineage tracking, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and creating runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations showcasing migration work and key learnings
- Gap analysis, suggested follow-up actions, and handover of training materials
- Reference materials, further learning paths, and support options
Requirements
- A solid grasp of data engineering fundamentals
- Proficiency with SQL and stored procedures (e.g., Synapse / SQL Server)
- Knowledge of ETL orchestration concepts (such as ADF or similar tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers moving from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing the adoption of Databricks
Testimonials (1)
All the topics covered, although many were very quick, give us an idea of what we will need to delve into further. Additionally, I liked that we got to do some hands-on practice, although I still believe the course deserves more.
Sandra Mariela Lopez Bernal - Kueski
Course - Databricks
Machine Translated