Infrastructure Engineer
Infrastructure Engineer owns system stability, observability, and debugging at scale. Requires 3+ years experience with Go, Kubernetes, and tools like Datadog/Prometheus for production incident response.
About the job
Must-Have
- 3+ years of hands-on experience debugging production systems (logs, traces, incidents, etc.)
- Strong problem-solving skills and ability to dive into unfamiliar backend codebases
- Strong Go and Kubernetes experience
- Familiarity with observability and monitoring tools (e.g., Datadog, Prometheus, Sentry)
- Clear, calm communication under pressure — especially during live incidents
Nice-to-Have
- Experience working with distributed systems or services at scale
- Built or maintained internal tooling for on-call teams or reliability workflows
- Familiarity with deployment pipelines, CI/CD, or infra-as-code
- Experience improving system observability (e.g., custom metrics, traces, log pipelines)
Skills
Kubernetes, Go, Datadog, Prometheus, Sentry, Distributed Systems, CI/CD, Observability, Infrastructure As Code, Debugging
Similar jobs
DevOps / SRE jobsBuild and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.