Member of Technical Staff - Infrastructure
Build and operate Kubernetes-based infrastructure for secure, reliable enterprise AI deployments across cloud and customer-managed environments. The role requires production Kubernetes experience, cloud infrastructure expertise, infrastructure as code, networking, security, and deployment automation.
About the job
Responsibilities
- Design, build, and operate infrastructure for enterprise deployments, including single-tenant and customer-hosted environments.
- Build deployment systems for cloud-native and enterprise environments.
- Own Kubernetes-based infrastructure for services, sandboxed agent execution, compute workloads, storage mounts, networking, and observability.
- Support scientific and computational workloads, including HPC-adjacent workflows, bioinformatics tools, batch execution, GPU/CPU compute, and large-scale file and data movement.
- Work with customer-facing teams on enterprise deployment planning, debugging, reliability reviews, and production operations.
- Help define the long-term architecture for deployments in pharma, biotech, academic, and research environments.
Requirements
- 2+ years of industry experience in infrastructure engineering, DevOps, SRE, platform engineering, cloud engineering, or a similar role.
- Strong hands-on experience with Kubernetes in production.
- Experience deploying and operating systems on AWS, Google Cloud, or Azure.
- Experience with infrastructure-as-code tools such as Terraform, Pulumi, CDK, or similar.
- Strong understanding of networking, IAM, VPCs, load balancers, DNS, certificates, secrets management, and cloud security fundamentals.
- Experience building CI/CD pipelines, deployment automation, environment provisioning, and operational tooling.
Nice-to-Haves
- Experience deploying enterprise SaaS into single-tenant, customer-hosted, VPC-isolated, or regulated environments.
- Experience building or maintaining R&D infrastructure in pharma, biotech, academic research, or life sciences organizations.
- Experience with HPC, bioinformatics infrastructure, scientific computing, workflow engines, or batch schedulers.
- Experience with Nextflow, Snakemake, Cromwell, Slurm, AWS Batch, EKS, ECS, FSx, Lustre, S3, or distributed filesystems.
- Experience with secure sandboxing or isolated execution environments such as gVisor, Kata Containers, Firecracker, or Kubernetes sandboxing.
- Experience supporting GPU-enabled workloads, large-scale data processing, or scientific compute clusters.
- Experience with enterprise identity, SSO, audit logging, compliance controls, customer networking, VPNs, private connectivity, or data residency requirements.
- Experience working directly with enterprise customers during deployment, onboarding, debugging, or production support.
Compensation and Benefits
- Competitive salary and equity share.
- Full medical, dental, and vision coverage, including free therapy sessions and an eyewear stipend.
- 401(k) (US only).
- Unlimited PTO (US only).
- Lunch and snacks in the office.
- Regular team offsites and company events.
Skills
Kubernetes, AWS, GCP, Azure, Terraform, Pulumi, Aws Cdk, Networking, IAM, Vpc, CI/CD, Python, Slurm, S3
Similar jobs
DevOps / SRE jobsBackend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.
Production Engineer builds and operates large-scale systems, focusing on automation, monitoring, infrastructure management, and resilient operations. Requires 2+ years in SRE/DevOps, expertise in Linux, AWS, Kubernetes, and programming in Python or Golang.
Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.
Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.
Build and operate robust infrastructure, support enterprise deployments, and improve on-premises delivery for a rapidly scaling AI code review platform. The role requires networking expertise, cloud and container experience, and at least one year of infrastructure or software engineering experience.