Databricks Data Engineer
Melbourne, Victoria, Australia · Contract
Be the first to apply
- Experience
- 5–8 yrs
- Salary
- —
- Openings
- 1
- Posted
- 43 minutes ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
Viable Solutions is looking for a skilled Databricks Engineer based in Melbourne, VIC to join on a contract basis. The role demands extensive experience in constructing scalable, high-performance data engineering and analytics solutions for enterprise-level clients, leveraging the Databricks environment. The successful candidate will design, develop, and fine-tune data pipelines and implement lakehouse architectures with a focus on large-scale data operations and cloud-native systems.
Core Duties
- Architect and develop data pipelines and workflows on the Databricks Lakehouse Platform.
- Optimize and create Apache Spark jobs to handle extensive data processing and transformation tasks.
- Manage Delta Lake tables, ensuring schema management, performance optimization, and time travel functionality.
- Apply Databricks features such as Workflows and Delta Live Tables (DLT) to orchestrate pipelines effectively.
- Configure and optimize Databricks clusters, including autoscaling and cost efficiency strategies.
- Develop reusable notebooks and libraries using Python, Scala, or SQL languages.
- Construct ELT/ETL pipelines for data ingestion, transformation, and loading from diverse data formats including structured, semi-structured, and unstructured sources.
- Implement lakehouse architectural patterns across Bronze, Silver, and Gold data layers.
- Integrate Databricks systems with upstream and downstream components such as databases, APIs, and data warehouses.
- Establish data quality, validation, and monitoring measures throughout pipelines.
- Manage Unity Catalog for governance, lineage tracking, and access control.
- Lead the deployment and management of MLflow experiments, aiding in model tracking and registry processes.
- Support operationalization of machine learning models and develop feature engineering pipelines for ML workloads.
- Deploy Databricks platforms on cloud services including AWS, Azure, or GCP, and establish infrastructure as code using Terraform or Databricks Asset Bundles.
- Design and maintain CI/CD automation pipelines leveraging tools like GitHub Actions, Azure DevOps, or Jenkins, with adoption of GitOps for source control.
- Monitor system workloads for performance enhancements and cost reductions.
- Ensure robust data governance and security policies through Unity Catalog, implementing row-level and column-level security measures.
- Document data models, architectural designs, and operational procedures comprehensively.
Qualifications and Expertise
- Minimum 5 to 8 years of professional experience in data engineering or closely related fields.
- Proven mastery of the Databricks Lakehouse Platform is essential.
- Strong command over Apache Spark technologies including PySpark, Spark SQL, and Spark Structured Streaming.
- Advanced Python abilities focusing on data engineering and pipeline coding.
- Experience with Delta Lake functionalities, including table management, optimization, ACID transactions, and time travel.
- Familiarity with Delta Live Tables and Databricks Workflows for pipeline orchestration.
- Competence in complex SQL querying and managing data transformations.
- Hands-on knowledge of Unity Catalog for governance and controlled data access.
- Experience with MLflow for experiment and model management.
- Practical experience on cloud platforms—AWS, Azure, or Google Cloud Platform.
- Proficiency in continuous integration/deployment environments and tools such as GitHub Actions, Azure DevOps, or Jenkins.
- Expertise implementing infrastructure as code primarily using Terraform.
- Experience working within Agile or Scrum development methodologies.