Staff Site Reliability Engineer
Dublin, County Dublin, Ireland · Full Time
Be the first to apply
- Experience
- 8–10 yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Company
Replit is an agentic software platform that empowers anyone to create applications through natural language. Serving millions globally, it aims to democratize software development by removing traditional obstacles to app building.
About the Role
We are looking for a Staff Site Reliability Engineer to join our SRE team in Dublin. Your core responsibilities include maintaining and enhancing the performance, scalability, and reliability of Replit's infrastructure that supports millions of developers worldwide. You'll bridge engineering and operations, automate processes, and champion best practices for scalable, highly available systems.
Key Responsibilities
- Design, implement, and lead the deployment of observability tools including monitoring, logging, and tracing systems, creating dashboards that provide real-time insights into system health.
- Collaborate with product and engineering teams to define, execute, and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs), ensuring balance between reliability and innovation pace.
- Lead incident response initiatives; oversee high-impact incidents with a calm, strategic approach, conduct blameless post-mortems, and implement preventive mechanisms.
- Develop and refine automation and infrastructure as code frameworks such as Terraform or Pulumi to reduce manual toil and enable self-healing systems.
- Work with core infrastructure and product teams to fine-tune Kubernetes, Docker, and GCP cloud deployments, enhance capacity planning, and reduce latency globally.
- Diagnose complex issues across distributed systems, design durable solutions, and improve system operability and debug-ability.
- Provide senior-level design reviews focusing on reliability, scalability, security, and operational robustness.
- Mentor engineers across skill levels to instill reliability as a foundational culture attribute.
- Write clean, tested Python or Go code to develop internal tools and integrate third-party systems.
Required Qualifications
- Between 8 to 10 years’ experience in Site Reliability Engineering, DevOps, Systems or Infrastructure Engineering.
- Strong programming proficiency in Python or Go, with ability to deliver well-tested, high-quality code.
- Expertise in distributed systems design, service-oriented architecture, production service scaling, and maintenance.
- In-depth experience with container orchestration (Kubernetes) and cloud-native technologies.
- Proven ability to create and maintain complex monitoring and observability systems involving logs, metrics, and traces.
- Advanced skills in managing and resolving incidents under pressure and conducting rigorous incident reviews.
- Familiarity with infrastructure as code tools like Terraform or Pulumi and configuration management.
- Excellent communication skills capable of articulating technical concepts clearly and embracing transparency.
- Strong interpersonal skills with a history of mentoring engineers from junior to principal levels.
- A proactive attitude to understanding and improving all parts of the technology stack.
- Passion for democratizing software creation and enabling future generations of creators.
Preferred Bonuses
- Extensive knowledge of Google Cloud Platform services and tools.
- Expertise with observability platforms such as Prometheus, Grafana, Datadog, and OpenTelemetry.
- Experience crafting systems supporting high throughput and low latency.
- Deep hands-on experience with Go programming and Terraform scripting.
- Exposure to fast-growth startup environments.
- Experience in producing technical blog content and educational materials.
Benefits
- Competitive salary and equity compensation.
- 401(k) plan with 4% employer match (US-based employees only).
- Comprehensive health, dental, vision, and life insurance.
- Short-term and long-term disability coverage.
- Paid parental, medical, and caregiver leave.
- Flexible time off policy plus holidays.
- Commuter benefits for in-office and US employees.
- Monthly wellness stipend.
- Autonomous work environment.
- Reimbursement for in-office setup (in-office only).
- Quarterly team gatherings.
- On-site amenities for in-office staff.
Additional Information
Replit actively fosters an inclusive culture and encourages applications from candidates of diverse and underrepresented backgrounds to help achieve its mission of making programming universally accessible.
Level
Mid
Industry
Software Development