Skip to content
FloQastFloQast

Senior DevOps Engineer, AI Platform

Owns the multi-region AI runtime and platform reliability for customer-facing products, including Bedrock infrastructure, sandboxed execution, observability, cost controls, Terraform, and delivery pipelines. Requires 5+ years in DevOps, SRE, platform, or infrastructure engineering plus production experience operating LLM-backed workloads.

About the job

Responsibilities

  • Own AWS Bedrock and Bedrock AgentCore infrastructure across US, EU, and AU regions, including model access, throughput, cross-region inference, quotas, throttling, retries, budgeting, logging, and tracing.
  • Operate sandboxed execution environments for AI-generated code with session limits, network controls, tenant isolation, and least-privilege IAM.
  • Develop and review Terraform for multi-account, multi-region AWS infrastructure.
  • Extend Grafana observability for token consumption, model and regional latency, throttling, retries, tool-call failures, sandbox outcomes, generation success, and agent traces.
  • Define journey-based SLOs and runbooks for AI-specific failure modes.
  • Manage AI cost attribution across inference, serving, sandbox compute, and telemetry.
  • Build GitHub Actions CI/CD pipelines for Node.js/TypeScript and Python services in NX monorepos.
  • Gate, version, feature-flag, and roll back model and prompt changes using evaluation and regression suites.
  • Participate in DevOps on-call and plan capacity around peak accounting periods.
  • Maintain SOC 2 and ISO 27001/42001 controls, including audit logging, encryption, access control, patching, asset inventory, and tenant isolation.

Requirements

  • 5+ years of experience in DevOps, SRE, platform, or infrastructure engineering, including production on-call responsibility.
  • Hands-on AWS experience with ECS/Fargate, Lambda, SQS, S3, IAM, VPC, networking, and ALB/NLB.
  • Production-scale Terraform experience with modules, state management, multi-region, and multi-account deployments.
  • Experience with CI/CD, Docker, image supply chains, and scaling policies; GitHub Actions preferred.
  • Production operation of an LLM-backed or ML-serving workload, including tokens, latency, throttling, and cost.
  • Experience with observability tools and practices such as Grafana, Prometheus, OpenTelemetry, distributed tracing, and user-experience SLOs.
  • Working fluency in Python or TypeScript/Node.js, with the ability to read and debug the other.

Nice to have

  • Multi-region infrastructure under data-residency constraints.
  • AI cost and performance optimization, including token accounting, prompt caching, batching, model routing, and capacity right-sizing.
  • Sandboxed execution of untrusted or generated code.
  • Atmos or comparable Terraform orchestration and NX, Turborepo, or Bazel.
  • Progressive delivery with Harness or similar feature flags.
  • FinOps tooling such as CloudZero and per-tenant cost attribution.
  • MongoDB, PostgreSQL, Snowflake, or EMR/Spark.
  • SOC 2 or ISO 27001 audit evidence, cloud-cost reduction, or regulated SaaS experience.

Compensation and benefits

  • The role includes on-call responsibilities and work supporting customer-facing AI products across multiple regions.
  • A PhD or formal ML credential is not required.

Skills

AWS, Aws Bedrock, Aws Bedrock Agentcore, Terraform, Ecs/Fargate, AWS Lambda, Amazon Sqs, Amazon S3, IAM, Docker, GitHub Actions, Grafana, Prometheus, OpenTelemetry, Python

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Coinbase

Coinbase

United States

Senior Software Engineer, Core Infra Systems
$186k+/yrRemote5+ YOEDevOps / SRE

Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.

Shield AI

Shield AI

Seattle, WA
Senior Site Infrastructure Engineer
$110k+/yrOn-site5+ YOEDevOps / SRE

Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.