Staff Machine Learning Engineer
Toronto, Ontario, Canada (Hybrid) · Full Time
Be the first to apply
- Experience
- 7+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 5 days ago
- Work mode
- Hybrid
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About OpenTable
OpenTable, a leader in the restaurant reservation industry for over 25 years and part of Booking Holdings, Inc., empowers restaurants and diners worldwide. With over 70,000 restaurant partners and millions of diners, we enable seamless discovery and booking experiences. Our culture centers on hospitality and caring for others, supported by a global team dedicated to innovation.
Role Overview
We seek a Staff Machine Learning Engineer to shape technical strategies for deploying and operating production-scale machine learning systems. This engineering-focused position involves collaborating across teams to transform experimental models into dependable, monitored production services and establishing standards for ML system design and operation.
Key Initiatives
- Develop AI Agents to enhance restaurant search and discovery
- Deliver personalized recommendations to diners
- Build and serve high-throughput predictive models for marketplace optimization
- Integrate tools into agentic platforms using LLM tool calls and protocols
- Leverage multimodal understanding of text, images, and geospatial restaurant content
- Create AI-driven platforms providing restaurant partners with insights on business performance and diner demand
Responsibilities
- Define architectures, deployment methods, and monitoring strategies for machine learning models
- Develop production ML services that maintain high throughput, low latency, and observability from inception through operation
- Set engineering standards covering testing, continuous integration/deployment, alerting, rollback, and on-call responsibilities
- Collaborate closely with Product Managers and stakeholders to scope and prioritize cross-team ambiguous problems
Qualifications
- Minimum 7 years of professional software engineering experience, with significant time developing and maintaining production ML systems
- Broadened perspective from multiple organizations on best practices, architectures, and tradeoffs in ML production environments
- Practical experience with major cloud platforms (AWS, GCP, or Azure) for model serving including deployment, scaling, and observability tools
- Strong fundamentals in distributed systems, API and service design, concurrency, latency and throughput management, thorough testing, and operational ownership with on-call duties
- Advanced proficiency in Python and at least one statically typed language, preferably Java
- Proven expertise in training, deploying, and monitoring ML models at production levels
- Full MLOps lifecycle experience including model monitoring, drift detection, retraining procedures, version control, rollout and rollback protocols, and incident handling
- Demonstrated leadership through directing long-term projects, influencing engineering decisions beyond immediate teams, and collaborating with cross-functional stakeholders
Preferred Expertise
- Experience serving large language models (LLMs) at scale including inference infrastructure, GPU optimization, batching, caching, and latency/cost tradeoffs
- Deep knowledge of machine learning techniques in ranking, recommendations, classification, natural language processing, retrieval-augmented generation, and agentic system design
- Kubernetes production experience at significant scale
- Background in developing ETL jobs (especially Spark) or managing data warehouse infrastructures
- Familiarity with A/B testing methodologies and analysis
- Experience leading platform tool introductions and adoption within teams
Technology Stack
- Data pipelines: Spark, Airflow, EMR, SageMaker, Snowflake, S3, Delta Lake
- Machine learning frameworks and tools: PyTorch, XGBoost, CatBoost, LLMs, LangChain, LangSmith
- Deployment and orchestration: Docker, Kubernetes, Helm, Prometheus, Graphite, Grafana
- Infrastructure components: Kafka, Elasticsearch, Postgres, MongoDB, Redis, Qdrant
- Build and development tools: Poetry, FastAPI, Flask, Gunicorn/Uvicorn, Spring, Maven, TeamCity
Work Environment & Culture
This role requires working onsite in Toronto two days per week, offering a blend of in-person collaboration and flexible remote work. Our ML team balances working with an extensive historical dataset and a lean team environment, necessitating critical thinking and prioritization skills.
Benefits and Perks
- Up to 20 days per year working from virtually anywhere
- Comprehensive mental health support including company-paid therapy (SpringHealth) and Headspace subscription
- Company-wide annual week off dedicated to full team recharge
- Paid parental leave and generous vacation including birthday leave
- Paid volunteer time
- Dedicated career growth programs including Development Dollars, leadership training, and access to thousands of e-learning resources
- Employee discounts on travel
- Inclusion in Employee Resource Groups
- Private health and dental coverage
- Life and disability insurance
Additional Information
OpenTable fosters a diverse and inclusive environment where every employee is valued and supported. Candidates requiring accommodations during application or employment may contact the recruitment team for support. Compensation considers market data, location, and candidate experience, with offerings of competitive salary, benefits, bonuses, and equity grants. The role involves occasional communications outside core business hours to collaborate globally while respecting local labor laws and regulations.
Level
Mid