Skip to content

Platform Operations Engineer

Builds and scales platform infrastructure on AWS EKS with GitOps via ArgoCD, manages CI/CD with GitHub Actions, drives observability using Datadog/Sentry/CloudWatch, and ensures reliability through SLOs and incident response. Requires 3+ years SRE/DevOps experience and Kubernetes expertise.

About the job

Responsibilities

  • Support the Platform Infrastructure - Help manage and scale our container environment on Amazon EKS, implement GitOps workflows using ArgoCD, and maintain CI/CD pipelines through GitHub Actions to ensure that deployments are fast, consistent, and automated
  • Build for Reliability - Define and track SLIs and SLOs, lead incident response including on-call rotations, root cause analysis, and post-mortems, and contribute to disaster recovery planning to keep our systems highly available
  • Drive Observability - Design and maintain our monitoring and logging stack using Datadog, Sentry, and CloudWatch — giving engineering teams clear visibility into system health and performance before problems reach users
  • Shape the Platform's Future - Collaborate on architectural decisions, build internal tooling and self-service workflows that make the platform easier to operate, and contribute meaningfully to how we scale and evolve our infrastructure

Requirements

  • 3+ years in SRE, DevOps, or Cloud Infrastructure
  • Confident working with core AWS services (VPC, IAM, EKS, RDS) and a strong understanding of cloud networking and security best practices
  • Expert in using Infrastructure as code with Terraform, CloudFormation, or Crossplane
  • Proficient with GitHub and GitHub Actions as a core component of your CI/CD and automation pipelines - not just for source control
  • Experienced with running Kubernetes clusters in production and managing application deployments through GitOps workflows (ArgoCD/Flux) and Helm Charts
  • Proficient with observability tooling such as Datadog, Sentry, CloudWatch, Grafana to include building alerts, dashboards, and log pipelines
  • Experience writing solid Python scripts to glue systems together, automate infrastructure tasks, or handle custom workflows
  • Comfortable working independently in a remote setup, asking questions when needed, and keeping momentum without being micromanaged
  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience

Nice to Haves

  • Certifications: AWS, Kubernetes, Terraform or Python

Benefits

  • Competitive pay with equity options
  • Stellar health care plan options (Medical, Dental & Vision), with FSA, DCFSA, & HSA options
  • Company-sponsored disability & life insurance
  • Unlimited PTO
  • 401(k) + 4% Matching
  • Fully remote work + flexible working hours
  • $750 work-from-home setup budget
  • Paid biannual in-person company summits
  • Quarterly $150 co-hanging stipend to meet up with coworkers
  • Monthly $100 health and wellness benefit
  • Generous paid family leave
  • Annual $1,200 learning & development stipend

Skills

AWS, EKS, Kubernetes, Terraform, Argo CD, GitHub Actions, Datadog, Sentry, CloudWatch, Python, GitOps, Helm, Grafana

Cloudflare

Cloudflare

Austin, TX
Systems Engineer - Database Platform
$150k+/yrHybridDevOps / SRE

Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.

Fluidstack

Fluidstack

New York, NY
Infrastructure Deployment Engineer
$150k+/yrOn-site5+ YOEDevOps / SRE

Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.

Trexquant

Trexquant

New York, NY

Python Engineer - Trade Operations
$150k+/yrOn-site3+ YOEDevOps / SRE

The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Hebbia

Hebbia

New York, NY
Software Engineer, Infrastructure
$160k+/yrOn-site5+ YOEDevOps / SRE

Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.