Site Reliability Engineer
Site Reliability Engineer enhances system observability, reliability, and availability at a prediction markets platform. Builds automation, optimizes cloud infrastructure (Kubernetes, Docker, Terraform), debugs issues, and participates in on-call rotations. Requires 4+ years software engineering experience.
About the job
What You’ll Do
- Improve observability, reliability, and service availability by defining and measuring key metrics
- Build automation and systems that eliminate toil and reduce operational burden
- Collaborate with core infrastructure engineers to performance-tune and optimize cloud deployments (Docker, Terraform, Kubernetes, EC2, etc.)
- Partner with product teams to minimize service disruptions and automate incident response
- Identify and analyze reliability problems across the stack, designing and implementing software for significant, long-term improvements
- Mentor engineers and drive a culture where reliability is a core engineering value
- Write high-quality, well-tested code that supports internal and external customer needs
- Debug complex technical issues and improve system usability, operability, and diagnosability
- Review feature designs across the company and ensure security, safety, scalability, and architectural clarity
- Build and maintain integrations with third-party vendors
- Participate in on-call rotations to troubleshoot and resolve urgent issues
What You Bring
- 4+ years of software engineering experience
- Experience designing, building, scaling, and maintaining production services and service-oriented architectures
- Strong system design, coding, debugging, performance-tuning, and observability skills
- High-quality coding practices with strong testing discipline
- Excellent written and verbal communication; comfort working transparently across teams
- Strong interpersonal skills across junior-to-principal engineering levels
- Ability to think clearly under pressure and dive into any layer of the stack
- Passion for building an open financial system that connects the world
- Willingness to participate in on-call rotations and swiftly resolve issues
Bonus Points:
- Experience designing highly reliable, high-throughput, low-latency systems
- Experience with Datadog
- Experience with Rust, Go, and Terraform
- Experience with AWS, GCP, or Azure
- Experience operating in regulated environments
- Experience writing training materials or company-facing engineering content
NYC Pay Transparency
Salary: $100,000–$250,000 annually, plus equity and benefits.
Skills
Kubernetes, Docker, Terraform, Datadog, AWS, GCP, Azure, Rust, Go, EC2
Similar jobs
DevOps / SRE jobsThe DevOps Engineer will design and operate AWS and hybrid infrastructure, improve CI/CD reliability, and strengthen disaster recovery and business continuity. The role requires 5+ years of cloud infrastructure experience, strong AWS proficiency, and hands-on infrastructure-as-code expertise.
Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.