S

Site Reliability Engineer

SnazzyHR

Remote · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
1 day ago
Work mode
Work from home
Education
Bachelor's degree in Computer Science or Engineering or related field
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About SnazzyHR

SnazzyHR is an innovative and user-focused HR and recruitment platform designed to alleviate the complexities and expenses associated with traditional hiring and HR management. Tailored for startups, recruiters, and expanding organizations, it delivers a scalable recruitment system free from legacy burdens and costly fees. The platform aims to enhance HR functions by streamlining core processes and empowering users.

Role Overview

The Site Reliability Engineer will ensure the consistent performance, availability, and robustness of SnazzyHR's cloud-hosted platform. In this fully remote full-time role, you will engineer, deploy, and manage scalable infrastructure solutions; oversee system health monitoring; and address incidents to reduce downtime. Collaboration with development teams to enhance system durability, automation of deployments and workflows, as well as conducting root cause analyses for production issues, are key components of the role.

Key Duties

  • Design, implement, and maintain scalable cloud infrastructure supporting the platform.
  • Monitor systems to ensure optimal health and performance, responding swiftly to incidents.
  • Collaborate with software engineers to strengthen resilience and reliability.
  • Automate deployment and operations workflows for efficiency and repeatability.
  • Perform thorough root cause analysis on production failures to drive continuous improvement.
  • Develop tools and methodologies to enhance observability, capacity planning, and fault tolerance.
  • Create and update comprehensive documentation of operational procedures and standards.

Qualifications

  • Extensive expertise in Site Reliability Engineering with a focus on designing and scaling cloud-based systems.
  • Strong skills in system administration and troubleshooting, particularly within Linux/Unix environments handling production issues.
  • Practical software development experience, especially in backend or tooling languages such as Python, Go, Java, or similar.
  • Proficient with monitoring, logging, alerting tools, CI/CD processes, and configuration management practices.
  • Good understanding of networking principles, security protocols, and architectures for high availability.
  • Independent worker skilled at remote communication and effective documentation.
  • Bachelor’s degree in Computer Science, Engineering, or related discipline, or equivalent practical experience.
  • Familiarity with cloud platforms (AWS, GCP, Azure) and container technologies (Docker, Kubernetes) is a strong advantage.

Minimum education

Bachelor's Degree

How they work

Relationship Building required

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer