Roles & Responsibilities
- Role Overview
- Serves as the core technical authority for designing, modernizing, and optimizing the Databricks Lakehouse architecture, code refactoring, and data execution pipelines.
- Key Roles & Responsibilities
- Refactoring & Modernization: Refactor legacy Synapse T-SQL stored procedures, scripts, and views into native PySpark, Databricks SQL, or Delta Live Tables (DLT).
- Storage Optimization: Transition legacy Synapse distribution schemas (HASH, ROUND_ROBIN, REPLICATE) to Delta Lake Liquid Clustering (CLUSTER BY) for optimal query performance and reduced maintenance.
- Pipeline Design: Design efficient batch and streaming ingestion patterns leveraging Auto Loader and Delta Lake Lakehouse mechanisms.
- Performance Tuning: Execute comprehensive tuning using OPTIMIZE, Z-ORDER, compute cluster sizing, and Serverless warehouse setup.
Education & Qualifications
Bachelor's degree in Computer Science/IT/MSC/MBA. Any Related Field.