Skip to content
HappyRobotHappyRobot

Infrastructure Engineer

Infrastructure Engineer owns system stability, observability, and debugging at scale. Requires 3+ years experience with Go, Kubernetes, and tools like Datadog/Prometheus for production incident response.

About the job

Must-Have

  • 3+ years of hands-on experience debugging production systems (logs, traces, incidents, etc.)
  • Strong problem-solving skills and ability to dive into unfamiliar backend codebases
  • Strong Go and Kubernetes experience
  • Familiarity with observability and monitoring tools (e.g., Datadog, Prometheus, Sentry)
  • Clear, calm communication under pressure — especially during live incidents

Nice-to-Have

  • Experience working with distributed systems or services at scale
  • Built or maintained internal tooling for on-call teams or reliability workflows
  • Familiarity with deployment pipelines, CI/CD, or infra-as-code
  • Experience improving system observability (e.g., custom metrics, traces, log pipelines)

Skills

Kubernetes, Go, Datadog, Prometheus, Sentry, Distributed Systems, CI/CD, Observability, Infrastructure As Code, Debugging

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.