Member Of Technical Staff
Build and operate low-latency infrastructure and distributed backend systems powering a high-throughput search stack. The role requires strong cloud, Linux, automation, observability, and systems-language experience across AWS, Kubernetes, Rust, and Go.
About the job
Responsibilities
- Build and operate scalable, low-latency infrastructure for search serving and retrieval.
- Design and improve distributed backend systems that handle high request volumes.
- Debug and optimize Linux systems, containers, networking, and production services.
- Improve reliability through observability, capacity planning, performance profiling, and incident response.
- Design, deploy, and operate cloud-native systems, primarily on AWS and Kubernetes.
- Build internal tools and automation for development, testing, debugging, and infrastructure operations.
- Improve CI/CD pipelines, release processes, and production rollout safety.
- Contribute directly to product codebases across Rust, Go, and other systems languages.
- Own systems end to end, from architecture and implementation to deployment and production operation.
Requirements
- Strong experience with cloud infrastructure, distributed systems, and automation.
- Deep understanding of Linux internals, performance analysis, and production debugging.
- Experience building or operating latency-sensitive, high-throughput backend systems.
- Experience with containers, orchestration, observability, and infrastructure as code.
- Experience building or maintaining CI/CD systems and release tooling.
- Fluency in at least one systems language, such as Rust, Go, C++, or Java.
- Comfort working across infrastructure and application-level code.
- Strong ownership and the ability to operate effectively in a fast-moving environment.
Nice to Have
- Experience with search, retrieval, ranking, or other large-scale data-serving systems.
- Experience operating Rust services in production.
- Experience with AWS, Kubernetes, and modern observability systems.
- Experience with profiling, capacity planning, and performance optimization.
- Experience building AI-assisted developer or operational tooling.
Skills
Rust, Go, C++, Java, AWS, Kubernetes, Linux, Docker, Distributed Systems, Observability, Infrastructure As Code, CI/CD, Cloud Infrastructure, Performance Profiling, Networking
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.