Software Engineer, Reliability
Builds and maintains scalable, reliable infrastructure including testing tools, automation, and resource management platforms for AI systems. Collaborates cross-functionally to ensure high availability, performance, and fault tolerance in a fast-paced environment.
About the job
Responsibilities
- Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands.
- Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable.
- Build and maintain automation tools to streamline repetitive tasks and improve system reliability.
- Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources.
- Implement fault-tolerant and resilient design patterns to minimize service disruptions.
- Develop and maintain service level objectives (SLOs) and service level indicators (SLIs) to measure and ensure system reliability.
- Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
- Participate in an on-call rotation to respond to critical incidents and ensure 24/7 system availability.
Requirements
- Proven experience as an SWE focused on reliability or a similar role in a fast-paced, rapidly scaling company.
- Strong proficiency in cloud infrastructure.
- Proficiency in programming languages.
- Experience with containerization technologies and container orchestration platforms like Kubernetes.
- Knowledge of IaC tools such as Terraform or CloudFormation.
- Excellent problem-solving and troubleshooting skills.
- Strong communication and collaboration skills.
- Experience with observability tools such as DataDog, Prometheus, Grafana and Splunk.
- Experience with microservices architecture and service mesh technologies.
- Knowledge of security best practices in cloud environments.
Nice-to-Haves
- Track record of accelerating engineering reliability by empowering fellow engineers with excellent tooling and systems.
- Experience utilizing Infrastructure as Code (IaC) principles to automate infrastructure provisioning and configuration management.
- Experience collaborating with cross-functional teams to ensure reliability and scalability in design and development.
Skills
Kubernetes, Terraform, CloudFormation, Datadog, Prometheus, Grafana, Splunk, Infrastructure As Code, Microservices, Cloud Infrastructure
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.