- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Job Overview
We are seeking a skilled Data Engineer specializing in Apache Doris to join our team. The ideal candidate must have direct experience working with Apache Doris or similar massively parallel processing (MPP) OLAP engines and exhibit strong expertise in advanced SQL and data architectures.
Key Responsibilities and Expertise
- Utilize and manage Apache Doris or comparable MPP OLAP engines such as StarRocks, ClickHouse, and Greenplum in production environments.
- Apply advanced SQL techniques including complex analytical queries, window functions, common table expressions (CTEs), and perform query optimization to enhance performance.
- Demonstrate comprehensive understanding of Apache Doris architecture, including front-end and back-end components, its data models (Duplicate, Aggregate, Unique), partitioning schemes, bucketing, and tablet/replica management.
- Handle Apache Iceberg or similar open table formats like Delta Lake or Hudi, and work with lakehouse architectures or external catalog federation.
- Understand and be able to work with Snowflake and/or Apache Spark to interpret, comprehend, and migrate existing workloads effectively.
- Comprehend distributed or MPP query execution concepts such as join distribution, runtime filtering, memory management strategies, and resolving data skew issues.
- Have hands-on production experience with Trino or PrestoSQL/Presto, focusing on the distributed query execution including MPP architecture, join distribution, memory spilling, partition pruning, and predicate pushdown.
- Expertise with cloud-based object storage systems and columnar file formats such as Parquet and ORC is essential.
- Proficiency in at least one programming language among Python, Java, or Scala, for developing tooling, creating user-defined functions (UDFs), and automation.
- Familiarity with version control systems (Git) and continuous integration/continuous deployment (CI/CD) pipelines for managing data workflows.
Mandatory Skills
- Proven experience and knowledge of Apache Doris.
Preferred Skills
- Experience with Apache Iceberg or other lakehouse technologies.
- Advanced SQL capabilities.
- Skills in Trino/Presto query engines.
- Experience working with Snowflake.
- Competence in Apache Spark platform.