Staff Software Engineer: Platform SRE
Leads the design, operation, and evolution of Flexport’s cloud infrastructure, platform tooling, observability, and incident response systems. Requires 10+ years of software, SRE, or infrastructure engineering experience, deep AWS expertise, and strong Terraform and automation skills.
About the job
Responsibilities
- Design, build, and operate reliable, scalable infrastructure in AWS and other cloud platforms.
- Evolve Kubernetes, Helm charts, Argo deployment systems, and deployment guardrails.
- Improve observability using OpenMetrics and Datadog.
- Set standards for Infrastructure as Code using Terraform.
- Build and support CI/CD pipelines with BuildKit and GitHub Actions.
- Maintain build tooling and artifact repositories, including Gradle, Bazel, npm, pnpm, Bun, Go, Cargo, Artifactory, ECR, and GitHub.
- Improve incident response tooling and lead cross-functional platform changes.
- Identify infrastructure risks proactively and improve resilience, observability, efficiency, security, and operability.
Requirements
- 10+ years of software, site reliability, or infrastructure engineering experience, including substantial production infrastructure work.
- Deep cloud infrastructure knowledge and hands-on AWS experience.
- Strong Infrastructure as Code skills, preferably with Terraform, and experience creating safe, reusable infrastructure patterns.
- Strong software engineering skills and the ability to build automation and production tooling in a general-purpose programming language.
- Experience leading incident response, conducting blameless postmortems or COE processes, and implementing follow-up actions.
- Sound judgment regarding security, IAM, networking, and change management.
- Excellent written and verbal communication, including design and incident documentation and cross-team stakeholder alignment.
- Ability to work independently, identify high-leverage problems, and operate effectively in ambiguity.
Nice to Have
- Experience operating AWS at scale across multiple accounts, regions, and availability zones.
- Experience with GitOps and continuous delivery tools such as Argo CD.
- Experience with identity and access management at scale, including OIDC.
- Experience designing disaster recovery, business continuity, and resilience testing programs.
- Experience improving cloud cost efficiency and resource utilization.
Compensation and Benefits
- 25 vacation days based on full-time employment.
- Employer-paid collective health insurance, including basic and additional packages.
- Defined pension contribution scheme.
- Equity program.
- Catered meals, breakfast, snacks, and soft drinks at the office.
- Commuting costs covered for employees living outside Amsterdam.
- Employee Assistance Program.
- Parental leave benefit.
- Hybrid work with three office days per week.
Skills
AWS, Kubernetes, Helm, Argo Cd, Openmetrics, Datadog, Terraform, Buildkit, GitHub Actions, Gradle, Bazel, Go, Cargo, OIDC, IAM
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.