Member of Technical Staff
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
About the job
Responsibilities
- Own cross-cutting infrastructure problems spanning compute, storage, networking, data, deployment, and reliability.
- Eliminate bottlenecks across retrieval and serving paths, build shared abstractions across deployment environments, and resolve failures across platform boundaries.
- Design, build, and operate distributed infrastructure for consumer, AI, and enterprise workloads, from architecture through production operation.
- Build shared abstractions, automation, and tooling that make infrastructure easier and safer to use.
- Improve performance, availability, scalability, and cost-efficiency across online request traffic and background workloads.
- Debug complex production issues across service and infrastructure boundaries and implement durable architectural improvements.
- Set technical direction, lead high-impact cross-team programs, and establish technical standards.
Requirements
- 4+ years of professional software engineering experience building and operating production backend, platform, or distributed systems.
- Experience owning complex production systems end to end and delivering sustained technical impact across teams.
- Ability to set technical direction, lead through influence, and contribute through architecture, design reviews, and mentorship.
- Strong software engineering skills in Python or another systems/backend language such as Go, Rust, C++, or Java.
- Meaningful experience across at least two infrastructure domains, including cloud platforms, distributed systems, Kubernetes, storage, databases, networking, data systems, developer infrastructure, or production reliability.
- Ability to develop depth quickly in unfamiliar systems, reason across software and infrastructure layers, and drive production incidents from diagnosis through durable resolution.
Skills
Python, Go, Rust, C++, Java, Cloud Platforms, Distributed Systems, Kubernetes, Storage, Databases, Networking, Data Systems, Developer Infrastructure, Production Reliability, Automation
Similar jobs
DevOps / SRE jobsBuild and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.