Staff Infrastructure Engineer
Build and operate Kubernetes-based infrastructure, cloud systems, CI/CD, observability, and reliability tooling for a high-scale prediction market platform. The role requires 8+ years of infrastructure or platform engineering experience and strong production expertise with Kubernetes, cloud providers, infrastructure as code, and software development.
About the job
Responsibilities
- Design, build, and operate Polymarket’s Kubernetes-based platform, including cluster architecture, workload scheduling, and multi-environment management.
- Build systems for autoscaling, load balancing, service discovery, and disaster recovery.
- Write production-grade code and internal tooling to manage infrastructure as code and eliminate operational toil.
- Build and maintain reliable CI/CD pipelines.
- Develop monitoring, logging, and alerting systems; participate in on-call rotations and lead root-cause analysis for production incidents.
- Architect and optimize AWS or GCP infrastructure for cost, performance, and security.
- Collaborate with backend, exchange, and product engineering teams on platform performance and growth.
Requirements
- 8+ years of professional infrastructure, platform, or DevOps engineering experience, ideally operating high-traffic production systems.
- Deep production experience with Kubernetes and Docker, including cluster operations, workload orchestration, and container lifecycle management.
- Strong software engineering fundamentals and ability to write production-grade code in Go, Python, or a similar language.
- Experience with AWS or GCP, including networking, IAM, compute, and storage.
- Experience with infrastructure-as-code tooling such as Terraform or Pulumi and CI/CD systems.
- Ability to debug complex distributed production systems on Linux and own correctness, performance, and uptime.
Nice-to-haves
- Experience operating high-throughput, low-latency, or financial/trading infrastructure.
- Experience with service mesh, GitOps tools such as ArgoCD or Flux, or multi-cluster/multi-region architectures.
- Familiarity with observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry.
Compensation and Benefits
- Competitive salary and equity.
- Unlimited PTO.
- Full health, vision, and dental coverage.
- 401(k) match.
- Hardware setup including a new MacBook Pro, large display, and accessories.
Skills
Kubernetes, Docker, Go, Python, AWS, GCP, Terraform, Pulumi, CI/CD, Linux, Prometheus, Grafana, Datadog, OpenTelemetry, Argo CD
Similar jobs
DevOps / SRE jobsStaff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.
Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.
Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.