Platform Engineer
Senior Platform Engineer owning AWS/EKS infrastructure, IaC with Terraform, GitOps, security, observability, and data systems for a fast-growing expert network marketplace. Requires 5+ years production Kubernetes/AWS experience, strong IaC and security skills.
About the job
What you'll do
- Run and scale production AWS and EKS, including cluster lifecycle, autoscaling, and platform add-ons (ingress, DNS, certificates, secrets).
- Manage cloud and cluster configuration declaratively with Terraform, Helm, Kustomize, eksctl and GitOps.
- Ship through GitHub Actions and GitOps (Argo CD / Flux), with automated build, test, and dependency gating.
- Handle SAST/DAST scanning, secrets management, and least-privilege IAM (IRSA), plus routine dependency and vulnerability remediation.
- Own metrics, logs, and traces across Datadog and Sentry.
- Operate PostgreSQL, MongoDB / Atlas, OpenSearch / Elasticsearch, and Temporal.
- Build automation and internal tooling in Python, Bash, Go, and Node.js / TypeScript; containerize with Docker.
What you bring
- 5+ years running AWS and Kubernetes in production (EKS preferred), including platform add-ons and AWS-native integration.
- Strong with Terraform and Kubernetes manifests (Helm, Kustomize), plus GitOps delivery (Argo CD and/or Flux).
- Security mindset: SAST/DAST, secrets management, least-privilege IAM, and vulnerability management.
- Observability with metrics, logs, and traces using Datadog, Sentry, or equivalents.
- Programming experience with Node.js / TypeScript and Python, plus Bash scripting on Linux.
- Experience with PostgreSQL and MongoDB; messaging/workflows with Temporal.
- Git / GitHub workflows and a documentation-first, async communication style.
- Networking fundamentals: TCP/IP, DNS, routing, load balancing.
Bonus experience
- AI / ML infrastructure (model-training pipelines, vector search).
- Cloud cost and capacity management at scale.
- Compliance-sensitive environments (e.g., SOC 2, ISO 27001).
Compensation
Full-time offers include base salary, equity, and benefits. Pay range: $160,000-$180,000 based on seniority, relevant experience and location.
Skills
AWS, Kubernetes, EKS, Terraform, Helm, Kustomize, Argo Cd, Flux, GitHub Actions, Datadog, Sentry, Postgres, MongoDB, Python, TypeScript
Similar jobs
DevOps / SRE jobsBuild and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.