AI Platform Engineer
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
About the job
Responsibilities
- Design, build, and maintain production-grade AI agent infrastructure across multiple cloud providers.
- Own the Azure or Oracle Cloud variant of the infrastructure and support Google Cloud and AWS environments.
- Enable portable infrastructure across customer environments.
- Improve observability, security, and system reliability.
- Deploy and operate applications using Kubernetes and Terraform.
- Partner with product and platform engineers to unblock delivery.
- Design and build CI/CD pipelines and own automation for infrastructure and application build, test, and deployment workflows.
- Build and maintain reusable automation frameworks for environment provisioning, secret rotation, dependency updates, compliance checks, release management, and developer workflows.
- Automate operational runbooks and recurring processes through scripts, pipeline stages, and self-service tooling.
- Promote observable, maintainable, automation-first engineering practices.
Requirements
- 5+ years of experience as a cloud infrastructure engineer.
- Strong hands-on experience with Terraform and Kubernetes.
- Strong hands-on experience with at least one of Azure, Oracle Cloud, Google Cloud, or AWS.
- Experience designing and owning CI/CD systems, pipeline architecture, GitOps deployments, release automation, and pipeline-as-code practices.
- Experience building DevOps and developer-workflow automation frameworks using Bash, Python, or similar scripting languages, workflow orchestration, and self-service tooling.
- Strong networking fundamentals.
- Proven experience deploying and operating production systems.
Nice-to-Haves
- Experience with FluxCD or other GitOps controllers such as Argo CD.
- Understanding of cloud security and DevSecOps, including security gates, SBOM generation, and compliance checks in CI/CD pipelines.
- Deep networking expertise.
- Experience with software development and the software development life cycle.
- Startup or big-tech experience.
- Strong ownership mindset and interest in building reliable foundational systems.
Skills
Terraform, Kubernetes, Azure, Oracle Cloud, GCP, AWS, CI/CD, GitOps, Fluxcd, Argo Cd, Bash, Python, DevSecOps, Sbom, Networking
Similar jobs
DevOps / SRE jobsElectrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.