- Experience
- 4+ yrs
- Salary
- USD 130,000 – USD 200,000 / year
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We are seeking a Site Reliability Engineer to ensure our production systems remain reliable, easily observable, and performant. This position blends software engineering responsibilities with operational tasks, focusing on automating repetitive work and prioritizing system reliability as a critical engineering challenge.
Key Responsibilities
- Design and manage systems to achieve high availability and optimal performance.
- Create and sustain observability tools, including logging, metrics collection, and tracing.
- Establish, monitor, and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets.
- Lead incident management including coordinating response efforts and conducting thorough post-mortem analyses.
- Automate repetitive operational tasks through developing tools and advancing platform capabilities.
- Collaborate with application development teams to ensure production readiness.
Required Skills and Qualifications
- Minimum of four years’ experience in site reliability engineering, DevOps, or infrastructure engineering roles.
- Proficient in scripting and software development, preferably using Python, Go, or similar languages.
- Extensive experience with leading cloud platforms such as AWS, Google Cloud Platform, and Microsoft Azure.
- Hands-on expertise with Kubernetes, Terraform, and observability solutions.
- Proven track record leading incident response in live production systems.
- Solid understanding of distributed systems architecture and concepts.
Additional Attributes
- Curiosity and initiative to analyze systems deeply and convert insights into tangible improvements.
- Strong written communication skills with the ability to articulate technical choices clearly.
- Adaptable mindset focused on rapid iteration through testing and learning cycles.
- Comfortable working asynchronously across multiple time zones.
Compensation and Benefits
This full-time remote role offers a competitive salary ranging between $130,000 and $200,000 (USD equivalent), depending on experience and qualifications, complemented by benefits that vary by location. Additional perks include a performance-based bonus, an annual stipend for learning and development, flexible work hours, and health and wellness programs. The position involves contributing to impactful projects at scale.
Equal Opportunity
We value diversity and are an equal opportunity employer. We encourage applications from all qualified candidates regardless of race, ethnicity, gender identity or expression, sexual orientation, religion, disability, age, or any other legally protected factor. All hiring decisions focus exclusively on the candidate's skills, qualifications, and ability to perform the job.