- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 4 days ago
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science or related discipline
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
We seek a seasoned Data & AI Engineer to enhance our team’s efforts across enterprise and public-sector projects focused on digital transformation, data platforms, cloud services, and artificial intelligence solutions.
Key Responsibilities
- Construct and sustain both batch and streaming data ingestion and transformation workflows.
- Architect and implement lakehouse frameworks, refined data schemas, data products, and APIs that serve data efficiently.
- Establish mechanisms for managing data quality, metadata, lineage tracking, classification, and access controls.
- Develop data pipelines that back machine learning, AI applications, document analytics, embedding generation, vector search, and Retrieval-Augmented Generation (RAG) systems.
- Deliver modular, maintainable, and scalable software solutions primarily using Python, SQL, and Apache Spark.
- Adopt rigorous software engineering practices such as automated unit testing, code reviews, continuous integration and deployment.
- Engage with cloud infrastructure, containerization, orchestration platforms, and monitoring/observability tools.
- Diagnose and resolve data or application-related production issues and ensure smooth operational support.
- Leverage enterprise-sanctioned AI coding assistants like GitHub Copilot or Microsoft Copilot for coding, testing, documentation, and analysis tasks.
- Critically validate AI-generated code for accuracy, security, compliance with licensing, efficiency, and maintainability independently.
- Collaborate effectively with data scientists, engineers, architects, business stakeholders, and project teams.
Qualifications and Skills
- Minimum of five years of professional background in data engineering, AI engineering, software development, or related roles.
- Advanced proficiency in Python programming and SQL querying.
- Demonstrated experience designing and implementing ETL/ELT data pipelines and processing solutions.
- Deep familiarity with Apache Spark and contemporary data/lakehouse architectures.
- Hands-on experience with both batch and real-time streaming data processing frameworks.
- Competence with API development, version control systems (Git), automated testing, and CI/CD pipelines.
- Working knowledge of cloud data platforms, container technologies, orchestration tools, and observability monitoring.
- Understanding of machine learning data workflows including preparation, embedding techniques, vector databases/search, and RAG technology.
- Comprehensive grasp of data governance fundamentals such as quality control, lineage, metadata management, and access restrictions.
- Experience operating in complex enterprise or government environments is highly valued.
- Proven ability to independently review and verify AI-assisted code for quality and compliance.
Educational Background and Certifications
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related area.
- Certification preferences include: Databricks Certified Data Engineer Associate, Microsoft Certified: Azure Data Engineer Associate, AWS Certified Data Engineer – Associate, Google Cloud Professional Data Engineer, and Microsoft Certified: Azure AI Engineer Associate.
Additional Technical Expertise (Preferred)
- Experience with Databricks, Azure, AWS, GCP cloud environments.
- Familiarity with Docker, Kubernetes, Kafka, Airflow workflow orchestration.
- Knowledge of lakehouse design and vector database implementations.
- Exposure to large language models (LLMs), GitHub Copilot, Microsoft Copilot AI coding assistants.
Minimum education
Bachelor's Degree
Skills
Data Engineering
Data Governance
Python Programming
· 5 to 8 years
Automated Testing
SQL Querying
· 5 to 8 years
Machine Learning Workflows
Version control with Git
Continuous Integration/Continuous Deployment
data lakehouse architecture
Cloud platforms (AWS)
API design and development
Containerization and orchestration
Tools & software
Apache Spark
· 5 to 8 years required