Senior DevOps Engineer, AI Platform
Owns the multi-region AI runtime and platform reliability for customer-facing products, including Bedrock infrastructure, sandboxed execution, observability, cost controls, Terraform, and delivery pipelines. Requires 5+ years in DevOps, SRE, platform, or infrastructure engineering plus production experience operating LLM-backed workloads.
About the job
Responsibilities
- Own AWS Bedrock and Bedrock AgentCore infrastructure across US, EU, and AU regions, including model access, throughput, cross-region inference, quotas, throttling, retries, budgeting, logging, and tracing.
- Operate sandboxed execution environments for AI-generated code with session limits, network controls, tenant isolation, and least-privilege IAM.
- Develop and review Terraform for multi-account, multi-region AWS infrastructure.
- Extend Grafana observability for token consumption, model and regional latency, throttling, retries, tool-call failures, sandbox outcomes, generation success, and agent traces.
- Define journey-based SLOs and runbooks for AI-specific failure modes.
- Manage AI cost attribution across inference, serving, sandbox compute, and telemetry.
- Build GitHub Actions CI/CD pipelines for Node.js/TypeScript and Python services in NX monorepos.
- Gate, version, feature-flag, and roll back model and prompt changes using evaluation and regression suites.
- Participate in DevOps on-call and plan capacity around peak accounting periods.
- Maintain SOC 2 and ISO 27001/42001 controls, including audit logging, encryption, access control, patching, asset inventory, and tenant isolation.
Requirements
- 5+ years of experience in DevOps, SRE, platform, or infrastructure engineering, including production on-call responsibility.
- Hands-on AWS experience with ECS/Fargate, Lambda, SQS, S3, IAM, VPC, networking, and ALB/NLB.
- Production-scale Terraform experience with modules, state management, multi-region, and multi-account deployments.
- Experience with CI/CD, Docker, image supply chains, and scaling policies; GitHub Actions preferred.
- Production operation of an LLM-backed or ML-serving workload, including tokens, latency, throttling, and cost.
- Experience with observability tools and practices such as Grafana, Prometheus, OpenTelemetry, distributed tracing, and user-experience SLOs.
- Working fluency in Python or TypeScript/Node.js, with the ability to read and debug the other.
Nice to have
- Multi-region infrastructure under data-residency constraints.
- AI cost and performance optimization, including token accounting, prompt caching, batching, model routing, and capacity right-sizing.
- Sandboxed execution of untrusted or generated code.
- Atmos or comparable Terraform orchestration and NX, Turborepo, or Bazel.
- Progressive delivery with Harness or similar feature flags.
- FinOps tooling such as CloudZero and per-tenant cost attribution.
- MongoDB, PostgreSQL, Snowflake, or EMR/Spark.
- SOC 2 or ISO 27001 audit evidence, cloud-cost reduction, or regulated SaaS experience.
Compensation and benefits
- The role includes on-call responsibilities and work supporting customer-facing AI products across multiple regions.
- A PhD or formal ML credential is not required.
Skills
AWS, Aws Bedrock, Aws Bedrock Agentcore, Terraform, Ecs/Fargate, AWS Lambda, Amazon Sqs, Amazon S3, IAM, Docker, GitHub Actions, Grafana, Prometheus, OpenTelemetry, Python
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.