Software Engineer, DevOps
Designs and builds scalable infrastructure for AI products, focusing on cloud platforms, Kubernetes orchestration, CI/CD pipelines, and observability. Requires 3+ years in infrastructure engineering and bachelor's/master's in CS.
About the job
Responsibilities
- Partner with product teams to architect, design, and build the foundational infrastructure for our products.
- Design, develop, and deploy highly available and scalable Multi-tenant SaaS solutions on public cloud networks like AWS, Azure, and GCP. Leverage technologies such as Kubernetes, Helm, Terraform, and Istio to achieve infrastructure resilience.
- Drive the automation of infrastructure tasks, from provisioning to configuration management and deployment, utilizing tools like Terraform, Ansible, and Kubernetes.
- Collaborate closely with the software development team to refine CI/CD pipelines, e.g., using GitHub Actions and Cloud Build tools, enhance service interfaces, and improve the overall developer experience.
- Architect and implement advanced observability solutions using tools like Prometheus and Grafana. Ensure real-time alerting and error tracking with Sentry and Pagerduty to maintain system health and performance.
- Deploy comprehensive testing frameworks, including tools like Selenium for end-to-end testing. Ensure robust integration and system testing to maintain software quality.
- Regularly monitor system health, analyze performance metrics, and recommend enhancements. This includes optimizing database queries and ensuring peak database performance.
Nice to Have
- MLOps experience
- Experience with Postgres query optimization and related performance improvement techniques.
- Experience with event-driven data and machine learning infrastructure, including streaming pipelines, database systems, model training
- Experience with air-gapped cloud environments or private clouds
- Experience administering complex deployments on Azure, especially AKS
Qualifications
- Bachelor's or Master's degree in Computer Science or related field.
- 3+ years of experience in Infrastructure engineering, or a similar role
- Excellent problem-solving skills and the ability to work under pressure in a fast-paced environment.
- Ability to work independently and as part of a team
- Experience working with global teams
Compensation (California based candidates)
- Standard base salary: $135,000-$225,000 annually. Compensation offered will be determined by factors such as location, level, job-related knowledge, skills, and experience. Certain roles may be eligible for variable compensation, equity, and benefits.
Skills
Kubernetes, Terraform, AWS, Azure, GCP, Helm, Istio, Ansible, GitHub Actions, Prometheus, Grafana, Sentry, Pagerduty, Selenium, Postgres
Similar jobs
DevOps / SRE jobsBuild and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.
Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.
Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.