Member of Technical Staff - System Engineering
Builds and operates core distributed systems, infrastructure, and service architecture for Phylo's agentic AI platform, ensuring reliable scaled execution across cloud and enterprise environments. Requires 3+ years in backend/infrastructure engineering with Kubernetes and cloud expertise.
About the job
Responsibilities
- Design and build production systems that orchestrate agent execution and power AI-driven scientific workloads.
- Build and operate scalable, reliable infrastructure across cloud, hybrid, and on-prem enterprise environments.
- Develop systems for sandboxed execution, secure task isolation, and controlled compute environments.
- Design and implement security, access control, and compliance foundations suitable for enterprise deployments.
- Partner closely with ML and science teams to translate computational workflows into robust, production-grade distributed systems.
Requirements
- 3+ years of industry experience in backend, infrastructure, or distributed systems engineering.
- Strong proficiency in at least one programming language (Python, Go, Rust, or similar).
- Experience designing and operating distributed systems in production.
- Deep hands-on experience with containerization and Kubernetes.
- Experience with infrastructure-as-code tooling (Terraform, Pulumi, or equivalent).
- Experience operating systems on at least one major cloud provider (AWS, GCP, or Azure).
- Comfort owning systems end-to-end in fast-moving, high-autonomy environments.
Nice to Haves
- Experience building enterprise SaaS platform that supports single-tenant, customer hosted deployment patterns.
- Experience in building R&D infrastructure at Pharma/Biotech.
- Experience with job orchestration, task scheduling, or workflow engines.
- Experience with sandboxed or isolated execution frameworks (gVisor, Kata Containers, Firecracker).
- Familiarity with distributed storage, observability systems, or high-performance compute environments.
Compensation & Benefits
- Competitive salary and equity share.
- Full medical, dental, and vision coverage, including free therapy sessions and eyewear stipend.
- 401(k).
- Unlimited PTO (US only).
- Lunch and snacks when in office.
- Regular team offsites and company events.
Skills
Kubernetes, Terraform, Python, Go, Rust, AWS, GCP, Azure, Pulumi, Distributed Systems
Similar jobs
DevOps / SRE jobsBuild and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.