Member of Technical Staff - Reliability Engineering
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
About the job
Responsibilities
- Define reliability standards, including SLOs, error budgets, and production-readiness criteria, using real system telemetry.
- Own the reliability toolchain: logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
- Ensure customer-facing reliability issues have clear ownership and are resolved.
- Identify and address reliability failures at system boundaries, including retry amplification, non-composing timeouts, and unmapped dependencies.
- Coordinate live production incidents, run blameless postmortems, and track follow-up actions to completion.
- Automate repetitive operational work to reduce on-call toil.
- Partner with cloud infrastructure, inference and training, performance, product, and control-plane teams on capacity, multi-region risk, serving and training failures, zero-downtime rollouts, and customer-facing reliability.
Requirements
- 5+ years of experience with Linux internals, system performance troubleshooting, and networking fundamentals, including TCP/IP, HTTP, and gRPC.
- 5+ years of experience writing production-grade tools and systems code in Python, Go, C++, or Rust.
- Experience operating and debugging Kubernetes, Terraform, and Docker in high-throughput production environments.
- Experience with high-throughput control planes, microservices, or multi-region systems.
- Knowledge of fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
- Ability to influence teams and drive adoption of standards through credibility and useful tooling.
- Willingness to work across unfamiliar parts of the stack when problems cross system boundaries.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
Nice-to-Haves
- Experience with Prometheus, Grafana, OpenTelemetry, and actionable alerting.
- Exposure to GPUs, inference serving, or distributed training.
- Experience building agents or LLM-based tooling for investigation, triage, or automation.
- Contributions to infrastructure, systems, or machine-learning serving open-source projects.
- Comfort working pragmatically and collaboratively in a startup environment.
Benefits and Compensation
- Solve challenging AI infrastructure problems, including low-latency inference and scalable model serving.
- Work with emerging technology used by businesses and developers globally.
- High ownership and direct impact in a fast-growing engineering organization.
- Collaboration with experienced engineers and AI researchers.
Skills
Linux, TCP/IP, Http, gRPC, Python, Go, C++, Rust, Kubernetes, Terraform, Docker, Prometheus, Grafana, OpenTelemetry
Similar jobs
DevOps / SRE jobsBuild and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.