Skip to content
OtterOtter

Production Engineer

Production Engineer builds and operates large-scale systems, focusing on automation, monitoring, infrastructure management, and resilient operations. Requires 2+ years in SRE/DevOps, expertise in Linux, AWS, Kubernetes, and programming in Python or Golang.

About the job

The Opportunity

We are seeking a talented Engineer with extensive system knowledge who wants to be part of a team to build and operate large-scale systems that enables reliable and rapid deployment with effective monitoring and resilient operations.

Your Impact

  • Automation! Capability! Performance! Scale!
  • Have a strong influence in defining our engineering best practices and deployment process
  • Help automate the continuous integration and testing processes to enable and scale
  • Manage and maintain infrastructure
  • Own, design and implement monitoring systems such as Prometheus and Grafana
  • Optimize Linux systems for performance, reliability, and security
  • Own configuration management process(es) and build product features as appropriate
  • Investigate and dig into data to find the root of a problem and strategize with our engineers on solutions
  • Participate in on-call rotation

We're looking for someone who

  • You have 2+ years experience in SRE, Production Engineering and/or DevOps
  • Expert level experience architecting, developing, and troubleshooting large scale systems
  • Advanced level proficiency with one or more programming languages (i.e. Python, Golang)
  • Deep experience with data structures and Linux systems internals (e.g., filesystems, system calls) and administration
  • Extensive experience with CI/CD pipelines and infrastructure as code (i.e. Terraform, Ansible)
  • You have a strong familiarity with AWS services (i.e. ECS, S3, ALB, VPC)
  • You have knowledge in containers and orchestration using Kubernetes
  • Experience building production quality cloud infrastructure that enables reliable and rapid deployment of large-scale systems with effective monitoring and resilient operations
  • You thrive working in a fast paced, startup environment
  • You have a proven track record taking on projects from inception to launch
  • Bachelor's in Computer Science or Electrical Engineering (MS preferred)

Salary Range

$155,000 to $185,000 USD per year.

Skills

Python, Go, Linux, AWS, Kubernetes, Terraform, Ansible, Prometheus, Grafana, CI/CD

Airtable

Airtable

San Francisco, CA
Software Engineer, Infrastructure (2-8 YOE)
$148k+/yrHybrid2+ YOEDevOps / SRE

Backend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.

Greptile

Greptile

San Francisco, CA

Infrastructure Engineer
$190k+/yrOn-site1+ YOEDevOps / SRE

Build and operate robust infrastructure, support enterprise deployments, and improve on-premises delivery for a rapidly scaling AI code review platform. The role requires networking expertise, cloud and container experience, and at least one year of infrastructure or software engineering experience.

Mercury

Mercury

San Francisco, CA
Software Engineer - Infrastructure
$116k+/yrRemote2+ YOEDevOps / SRE

Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.

Fab2

Fab2

Austin, TX
Infrastructure Software Engineering Intern
$114k+/yrOn-siteDevOps / SRE

Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff, Systems Infrastructure
$200k+/yrOn-siteDevOps / SRE

Build and operate large-scale scheduling, storage, caching, and networking infrastructure for AI training and inference. The role targets PhD researchers graduating by December 2026 with systems research depth and strong programming and performance-measurement skills.