- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 seconds ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Level
Level is a company specializing in learning technology, focused on empowering students to develop meaningful academic and life skills with enthusiasm and confidence. We integrate validated curriculum frameworks with outstanding interactive design to create engaging practice environments that complement classroom instruction, supporting educators, schools, and parents effectively.
Job Summary
As a Senior Infrastructure Engineer in our Platform team, you will be responsible for designing, constructing, and managing the cloud infrastructure and developer platform that underpin every Level product. Your role includes comprehensive ownership of critical infrastructure components such as Kubernetes platforms, infrastructure-as-code, CI/CD pipelines using GitOps, networking, DNS, observability, and cloud security postures. The position involves collaborating with a small, senior-leaning team, and infrastructure choices you make will directly influence the reliability, performance, expense, and development productivity.
Responsibilities
- Design, build, and maintain secure, highly available AWS cloud infrastructure leveraging Terraform/OpenTofu and adopting a GitOps workflow via Atlantis.
- Manage Kubernetes (EKS) operations including autoscaling with Karpenter, upgrades, core add-ons, and Helm deployment using ArgoCD.
- Create and sustain GitHub Actions pipelines to enable rapid and safe deployment by platform and product teams, fostering self-service tooling.
- Oversee networking components such as ingress/egress with Traefik, implement service mesh and mTLS with Linkerd/Envoy, manage load balancers and edge TLS, and maintain DNS using Route 53 with Terraform.
- Develop observability systems employing OpenTelemetry and SigNoz, utilizing telemetry data to enhance reliability, performance, and cost management. Lead escalation during complex incidents and conduct thorough post-mortems.
- Implement cloud security best practices pertaining to identity, secrets management, and network segmentation with particular attention to student data privacy and compliance, utilizing tools like Security Hub, GuardDuty, Inspector, Snyk, and organizational guardrails including Control Tower and SCPs.
- Provide leadership by defining standards, mentoring fellow engineers, and continuously improving the platform infrastructure.
Requirements
- Minimum 5 years of experience managing large-scale cloud infrastructure, preferably with AWS.
- Extensive expertise in infrastructure as code using tools such as Terraform/OpenTofu, with knowledge of CloudFormation, Pulumi, or CDK considered relevant.
- Strong production experience with containerization and orchestration using Docker and Kubernetes (particularly EKS).
- Proficiency in scripting and automation using languages like Python, Go, or Bash.
- Firm understanding of cloud networking concepts including VPC, DNS, load balancing, ingress, firewalls/WAF, and VPNs, coupled with best practices in security.
- Demonstrated ability with CI/CD and GitOps workflows, preferably with GitHub Actions or equivalent tools.
- Experience with observability tools and metrics (logs, traces) to inform operational improvements.
- Proven track record of independently leading complex infrastructure projects from inception through completion.
- Excellent communication skills capable of engaging technical and non-technical stakeholders.
Preferred Qualifications
- Experience with ArgoCD, Atlantis, Linkerd/Envoy, Traefik, Karpenter, and Helm.
- Familiarity with observability tooling such as OpenTelemetry, SigNoz, Grafana, Datadog, or Prometheus.
- Experience in building or working on internal developer platforms, including Backstage.
- Exposure to AI/ML infrastructure, including GPU resource scheduling, model and agent hosting, and inference gateway management.
- Knowledge of Rust service CI/CD pipelines and CloudFront/CDN technologies.
- Relevant professional certifications including AWS Solutions Architect or DevOps Engineer – Professional.
- Background in distributed systems and contributions to open-source infrastructure projects.
Level
Senior