Skip to content
FlexportFlexport

Staff Software Engineer: Platform SRE

Leads the design, operation, and evolution of Flexport’s cloud infrastructure, platform tooling, observability, and incident response systems. Requires 10+ years of software, SRE, or infrastructure engineering experience, deep AWS expertise, and strong Terraform and automation skills.

About the job

Responsibilities

  • Design, build, and operate reliable, scalable infrastructure in AWS and other cloud platforms.
  • Evolve Kubernetes, Helm charts, Argo deployment systems, and deployment guardrails.
  • Improve observability using OpenMetrics and Datadog.
  • Set standards for Infrastructure as Code using Terraform.
  • Build and support CI/CD pipelines with BuildKit and GitHub Actions.
  • Maintain build tooling and artifact repositories, including Gradle, Bazel, npm, pnpm, Bun, Go, Cargo, Artifactory, ECR, and GitHub.
  • Improve incident response tooling and lead cross-functional platform changes.
  • Identify infrastructure risks proactively and improve resilience, observability, efficiency, security, and operability.

Requirements

  • 10+ years of software, site reliability, or infrastructure engineering experience, including substantial production infrastructure work.
  • Deep cloud infrastructure knowledge and hands-on AWS experience.
  • Strong Infrastructure as Code skills, preferably with Terraform, and experience creating safe, reusable infrastructure patterns.
  • Strong software engineering skills and the ability to build automation and production tooling in a general-purpose programming language.
  • Experience leading incident response, conducting blameless postmortems or COE processes, and implementing follow-up actions.
  • Sound judgment regarding security, IAM, networking, and change management.
  • Excellent written and verbal communication, including design and incident documentation and cross-team stakeholder alignment.
  • Ability to work independently, identify high-leverage problems, and operate effectively in ambiguity.

Nice to Have

  • Experience operating AWS at scale across multiple accounts, regions, and availability zones.
  • Experience with GitOps and continuous delivery tools such as Argo CD.
  • Experience with identity and access management at scale, including OIDC.
  • Experience designing disaster recovery, business continuity, and resilience testing programs.
  • Experience improving cloud cost efficiency and resource utilization.

Compensation and Benefits

  • 25 vacation days based on full-time employment.
  • Employer-paid collective health insurance, including basic and additional packages.
  • Defined pension contribution scheme.
  • Equity program.
  • Catered meals, breakfast, snacks, and soft drinks at the office.
  • Commuting costs covered for employees living outside Amsterdam.
  • Employee Assistance Program.
  • Parental leave benefit.
  • Hybrid work with three office days per week.

Skills

AWS, Kubernetes, Helm, Argo Cd, Openmetrics, Datadog, Terraform, Buildkit, GitHub Actions, Gradle, Bazel, Go, Cargo, OIDC, IAM

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Phantom

Phantom

Remote

Staff DevOps Engineer
No salary listedRemote8+ YOEDevOps / SRE

Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.