Software Engineer, Research Infrastructure
Build and scale reliable research infrastructure and distributed systems for evolving AI research workflows. The role independently leads complex technical projects, makes foundational architectural decisions, and partners with researchers and engineering teams.
About the job
Responsibilities
- Design, build, and scale infrastructure and systems supporting rapidly increasing usage and evolving workloads.
- Independently scope and lead complex, ambiguous, multi-month engineering projects from initial concept through production.
- Drive cross-organizational alignment on technical direction across multiple stakeholders and teams.
- Make architectural decisions shaping research infrastructure and tooling.
- Partner with researchers to understand workflows and anticipate changing requirements.
- Iterate quickly using pragmatic, first-principles solutions and fast feedback loops.
- Own reliability and scalability as system load, usage, and complexity increase.
- Establish technical standards and best practices; mentor engineers.
Requirements
- Experience designing, building, and operating large-scale distributed systems or production infrastructure.
- Track record of independently scoping and delivering complex, ambiguous, multi-month technical projects.
- Strong software engineering fundamentals and hands-on coding ability.
- Experience making architectural decisions adopted by other engineers and teams.
- Strong written and verbal communication skills, including cross-team alignment.
- Ability to work effectively in ambiguous, fast-changing environments.
- Bachelor's degree or equivalent combination of education, training, and experience in a relevant field.
Nice-to-haves
- Infrastructure or platform experience supporting research or machine learning workflows.
- Experience addressing reliability and architectural challenges in rapidly scaling systems.
- Experience with distributed systems, cloud infrastructure, and infrastructure-as-code.
- Familiarity with compute, tooling, and workflow needs of large-scale machine learning research.
- Experience in a startup or startup-like environment.
- Experience as a technical lead or mentor.
Compensation
- Annual salary: $405,000–$625,000 USD.
Skills
Distributed Systems, Cloud Infrastructure, Infrastructure-As-Code, Machine Learning, Research Infrastructure, Scalability, Reliability Engineering, Software Architecture, Production Systems, System Design, Software Engineering, Technical Leadership
Similar jobs
Backend Engineering jobsBuild and operate secure, scalable sandboxing infrastructure for untrusted, model-generated code and tool calls. The role requires backend programming, virtualization or isolation experience, and strong knowledge of Linux security primitives.
Build backend systems, data infrastructure, agentic workflows, and enterprise integrations that make AI useful and trustworthy for financial institutions. The role requires 5+ years of software engineering experience, backend-language proficiency, and experience with production data-intensive or distributed systems.
Build and operate Host Assurance services and host software that establish trust in bare-metal and virtual machine infrastructure through secure bootstrap, identity, attestation, and verification. The role requires production software engineering experience across reliable systems, platform or infrastructure security, and host-system boundaries.
Build backend infrastructure and customer-facing workflows that enable developers and AI agents to operate reliably in secure cloud environments. The role requires strong Go and production backend experience, distributed-systems expertise, and practical knowledge of cloud infrastructure, networking, and security.
Build and own Baseten’s identity and authorization platform, including fine-grained permissions, credential systems, and enterprise administration. The role requires backend systems experience, production authorization expertise, and the ability to operate secure multi-tenant systems at scale.