- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
Our client is an expanding AI SaaS company located in Berlin seeking a skilled ML Data Engineer. This role involves constructing data pipelines, datasets, and infrastructure essential for training, evaluating, and deploying machine learning models in production environments.
You will play a key role bridging Data Engineering and Machine Learning, guaranteeing reliable, reproducible, and production-quality data support across the model development lifecycle.
Key Responsibilities
- Design and maintain scalable data pipelines utilizing Python and SQL.
- Create distributed data processing workflows leveraging Apache Spark.
- Manage orchestration of training and data workflows employing Apache Airflow.
- Develop dependable datasets for training and evaluating ML models.
- Construct ingestion and transformation pipelines for both structured and unstructured data sources.
- Set up automated data quality checks and validation protocols.
- Use MLflow to track datasets, experiments, and model artifacts.
- Build and manage data workloads on AWS infrastructure.
- Enhance reproducibility within ML training and evaluation phases.
- Monitor metrics such as data freshness, completeness, consistency, and quality.
- Diagnose and resolve production data-related issues impacting ML systems.
- Optimize pipeline performance to handle increasing volumes of data efficiently.
- Work closely with ML Engineers and Data Scientists to transition models from experimental stages into production deployment.
Required Qualifications and Experience
- At least three years of experience in Data Engineering, ML Data Engineering, ML infrastructure, or a closely related discipline.
- Proficiency in Python programming.
- Strong command of SQL for data manipulation and querying.
- Hands-on expertise with Apache Spark for distributed data processing.
- Experience managing workflow orchestration using Apache Airflow.
- Familiarity with AWS cloud services supporting data workloads.
- Knowledge in tracking ML lifecycle components using MLflow.
- Expertise in implementing data quality assurance and validation techniques.
- Demonstrated history of building and maintaining production-grade data pipelines.
Preferred Additional Skills
- Experience with deep learning libraries such as PyTorch or TensorFlow.
- Familiarity with Databricks platform.
- Knowledge of real-time streaming tools like Kafka.
- Working experience with Snowflake data warehousing.
- Understanding of data modeling and transformation tools like dbt.
- Expertise in feature engineering and managing feature stores.
- Experience in data versioning methodologies.
- Proficiency with data validation frameworks such as Great Expectations or Soda.
- Competence in containerization technologies like Docker and orchestration platforms including Kubernetes.
- Experience with infrastructure-as-code tools such as Terraform.
- Knowledge of ML model training pipelines and data lineage monitoring.
- Familiarity with observability practices in data workflows.
Skills
Tools & software
Apache Spark
· 2 to 5 years required
Apache Airflow
· 2 to 5 years required
Mlflow
· 2 to 5 years required
Amazon Web Services AWS
required