Software Engineer, Infrastructure
Builds scalable data and ML infrastructure supporting multi-cloud and deployment models. Partners with founders and engineers on core platforms, tooling, and research features for reliable production systems.
About the job
What You'll Work On
- Design and build the development and production platforms that power our products, enabling reliability and security at scale
- Architect, build, and deploy our core infrastructure while supporting multiple cloud providers and various deployment models
- Accelerate company productivity by empowering your fellow engineers & teammates with excellent tooling and systems, providing a best-in-case experience
- Partner with researchers and engineers to bring new features and research capabilities to our customers
About You
- Have meaningful experience in spearheading and constructing large-scale infrastructure
- Proficiency in bash, Kubernetes, Python, and/or Terraform or similar technologies
- Have experience working with AWS, other cloud platforms such as Azure or GCP and/or on-prem environments
- Have expertise in debugging problems across the stack, such as networking issues, performance problems, hardware issues or memory leaks
- Take pride in building and operating scalable, reliable, secure systems
- Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed
- Own problems end-to-end and are willing to pick up whatever knowledge you're missing to get the job done
We would love it if you had
- Built out data infrastructure from, or nearly from, scratch at a fast-growing startup
- Experience building ML/DL infrastructure and/or data infrastructure that feeds into training large ML models
Compensation
- Base salary: $180,000 to $300,000
- Significant equity
- 100% covered health benefits (medical, vision, and dental)
- 401(k) with 4% company match
- Unlimited PTO
- Annual $2,000 wellness stipend
- Annual $1,000 learning stipend
- Daily lunches and snacks
- Relocation assistance
Skills
Kubernetes, Terraform, Python, Bash, AWS, Azure, GCP, ML Infrastructure, Data Infrastructure, Debugging
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.