Staff Engineer - Observability Platform
Dublin, County Dublin, Ireland · Full Time
Be the first to apply
- Experience
- 10+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are looking to hire a Staff Engineer to design, develop, and operate the internal and external Observability infrastructure for MongoDB’s platform. This critical Observability stack supports tens of thousands of customers who rely on it to monitor their database clusters and receive timely alerts to protect vital workloads.
The Collections team, a recent addition within MongoDB’s Observability & Adoption Organization, is dedicated to simplifying telemetry onboarding and collection throughout MongoDB. This team owns essential elements of the observability collection framework, including onboarding workflows, telemetry collection agents across data and control planes, and ingestion services for metrics, logs, and traces. These systems support MongoDB’s internal and customer-facing observability, enabling insights, recommendations, and alerting.
Our mission focuses on minimizing obstacles for teams implementing and enhancing Observability, collaborating closely with development groups to instrument services with shared best practices. We define telemetry conventions and construct collection and ingestion systems that prioritize stability, performance, security, documentation, and self-service capabilities. We collaborate with the Data Pipeline and Storage & Query teams to ensure end-to-end performance and stability of MongoDB’s observability stack. Joining this team means influencing the observability approach across MongoDB and significantly impacting both developer experience and platform reliability.
As MongoDB Atlas and its infrastructure rapidly expand, we face growing demand for high-cardinality observability data for various use cases. Our systems process tens of billions of metric time series alongside petabytes of logs, traces, and event data. Our technology stack includes VictoriaMetrics, Grafana, Splunk, Flink, WarpStream/Kafka, Java, Golang, Fluentbit, and OpenTelemetry. Besides owning key observability infrastructure components, you will collaborate with SWE, Product, and SRE teams to promote and implement best practices for service instrumentation. This role is highly collaborative and offers ownership over some of MongoDB’s most critical internal infrastructure.
Our team embraces a culture of inclusivity, diversity, and collaboration. We welcome technically adept leaders eager to apply systems expertise to build foundational database infrastructure. We seek candidates based in Dublin for our hybrid working setup.
Profile and Requirements
- Enthusiasm for solving complex, high-scale problems to a high standard
- A minimum of 10 years’ experience designing, programming, debugging, and tuning distributed or highly concurrent mission-critical software using Go, Java, C, or C++
- Experience managing latency-sensitive, high-throughput systems
- Strong understanding of systems fundamentals such as multi-threaded programming, performance profiling, and advanced programming techniques
- Familiarity with database internals or designing core data processing system components
- Preferred: experience with indexing or database performance optimization
- Knowledge of the observability ecosystem and industry best practices
- Excellent verbal and written communication skills with a passion for collaboration and mentoring peers
- Proven project management and time management skills, with ability to realistically assess complexity and effort
- Good grasp of information security management principles
Key Responsibilities
- Architect and build systems for MongoDB’s mission-critical observability platform, focusing on performance, scalability, cost-efficiency, and resilience
- Develop observability enhancements that empower engineers and customers to quickly identify root causes of production issues
- Manage production customer escalations and provide coaching for the team to handle such situations effectively
- Write high-quality, production-ready database code and mentor colleagues to improve code standards
- Own critical codebases maintained by the Observability Team, ensuring quality, security, durability, availability, and maintainability
- Investigate and resolve test failures and production bugs, implementing preventative measures for new code
- Analyze the performance effects of code changes to avoid regressions
- Participate in interviewing and hiring advanced software engineering candidates
- Stay current with advancements in database and observability technology from industry and academic research
- Lead large-scale development projects and manage associated timelines
- Collaborate with stakeholders and engineering teams on cross-functional initiatives
- Advise Product Management on technical direction, complexity, and project dependencies
- Work closely with Product and Engineering leadership to define product roadmaps
- Promote strong unit and integration test practices to validate system correctness
- Participate in customer support rotations and work with customers to resolve issues
Success Milestones
- Within one month: Understand the high-level architecture of MongoDB observability and fix some bugs
- Within three months: Contribute to projects planned for the next major MongoDB release and resolve customer/test-reported issues
- Within six months: Take active code review responsibilities and participate in designing new features
- Within twelve months: Lead feature development, mentor teammates, and influence the long-term technical roadmap of the Observability Team
About MongoDB
MongoDB empowers customers and employees to innovate rapidly by redefining the data platform for the AI era. With the leading unified data platform available globally across AWS, Google Cloud, and Azure, over 67,000 customers, including 75% of the Fortune 100, rely on MongoDB for critical applications.
We are committed to fostering an inclusive and supportive work environment, offering benefits such as employee affinity groups, fertility assistance, and generous parental leave. We support individuals with disabilities throughout the application and interview process.
MongoDB is an equal opportunity employer.
Level
Mid
Industry
Software Development