Software Engineer - DevOps and MLOps
Builds and maintains infrastructure for software development, ML models, and AI operations including CI/CD pipelines, cloud platforms, and automation tools. Requires BS/MS in CS, experience with containerization, IaC, and programming in Python/Go/Java.
About the job
Responsibilities
- Design, implement, and manage CI/CD pipelines to facilitate seamless code integration and deployment.
- Monitor and optimize system performance, availability, and security.
- Automate infrastructure orchestration and configuration management using tools such as Kubernetes, Ansible, and similar.
- Configure and maintain data infrastructure appliances.
- Troubleshoot and resolve issues related to applications, infrastructure, and deployments.
- Work closely with our development and AI teams to deliver solutions that increase efficiency and stability.
Must-have Qualifications
- BS or MS in software engineering, computer science, or a related field.
- Proven experience standing up a CI/CD system from scratch.
- Experience with multi-language build systems (e.g., Bazel, Bob).
- Proficiency with cloud platforms (e.g., AWS, Azure, GCP) and containerization technologies (e.g., Docker, Kubernetes).
- Experience with automation tools (e.g., Terraform, Ansible, GitHub Actions, Jenkins) and version control systems (e.g., Git).
- Strong programming skills in languages such as Python, Go, or Java.
- Self-starter attitude with strong ability to identify problems, prioritize them, then plan and execute working solutions.
- Enthusiasm for working in a fast paced startup environment and eagerness to support the team on a variety of topics.
Nice-to-have
- Experience with MLOps platforms (e.g., MLflow, Kubeflow, or SageMaker).
- Knowledge of big data technologies (e.g., Hadoop, Spark, or Kafka).
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
- Understanding of machine learning frameworks (e.g., TensorFlow, PyTorch, or Scikit-Learn).
- Experience with edge computing and IoT device management.
- Knowledge of security best practices and compliance standards in AI/ML environments.
- Proficiency in database management systems (e.g., PostgreSQL, MongoDB, or Cassandra).
- Experience with infrastructure-as-code tools (e.g., CloudFormation, Pulumi).
- Knowledge of GitOps practices and tools (e.g., ArgoCD, Flux).
Skills
Kubernetes, Docker, Terraform, Ansible, CI/CD, GitHub Actions, Jenkins, Python, Go, AWS, Azure, GCP, Bazel, MLflow, Kubeflow
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.