Infrastructure Engineer
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
About the job
What You’ll Work On
- Designing and maintaining core infrastructure across cloud environments.
- Building Infrastructure-as-Code workflows to automate deployments and scaling.
- Improving monitoring, logging, and alerting systems to maintain reliability.
- Managing CI/CD pipelines (Github, Spacelift) to ensure seamless deployments.
- Supporting disaster recovery planning and ensuring high availability of systems.
- Partnering with product and research teams to design architectures that scale with workload demands.
- Identifying and fixing performance bottlenecks in compute, storage, and networking.
What We’re Looking For
- Strong experience with cloud platforms (AWS).
- Proficiency with Infrastructure-as-Code tools (Terraform).
- Deep experience with containers and orchestration (Docker).
- Solid understanding of distributed architectures.
- Programming languages (Python, Go).
- Proven ability to ship reliable infrastructure in production environments.
- Growth mindset and eagerness to thrive in a hyper-growth, high-ownership environment.
Benefits
- Generous equity grant vested over 4 years
- A $20K relocation bonus (if moving to the Bay Area)
- A $10K housing bonus (if you live within 0.5 miles of our office)
- A $1K monthly stipend for meals
- Free Equinox membership
- Health insurance
Skills
AWS, Terraform, Docker, Python, Go, Kubernetes, CI/CD, GitHub, Spacelift, Infrastructure As Code
Similar jobs
DevOps / SRE jobsBuild and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.
Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.