Software Engineer, Infrastructure
Builds and scales core infrastructure using MicroVMs, Kubernetes, AWS, and GCP to support AI agent development sandboxes. Requires 2+ years experience in infrastructure engineering and systems programming.
About the job
Responsibilities
- Design, develop, and maintain core infrastructure components, including MicroVMs, Kubernetes clusters, container builds, and cloud-based services.
- Optimize infrastructure for performance, scalability, and cost-efficiency.
- Collaborate with product engineering teams to integrate infrastructure with Runloop's AI software engineering platform.
- Implement robust monitoring and alerting systems to ensure the health and availability of our infrastructure.
- Stay up-to-date with the latest advancements in cloud infrastructure and AI technologies.
Qualifications
- Bachelor's degree in Computer Science or a related field, or equivalent experience.
- 2-4+ years of experience in software engineering, with a focus on infrastructure and platform development.
- Strong experience with MicroVMs, Kubernetes, AWS, and GCP.
- Experience in container optimization with projects like overlaybd, buildkit, etc.
- Proficiency in one or more systems programming languages, such as Golang, Rust, or Java.
- Experience with infrastructure as code (IaC) tools, such as Terraform or Pulumi.
- Familiarity with DevOps practices and tools, such as CI/CD pipelines, configuration management and observability tools.
- Excellent problem-solving skills and a passion for building scalable, reliable systems.
Bonus Points
- Experience with AI/ML workloads and infrastructure.
- Experience in container optimization with projects like discoball.
- Contributions to open-source projects.
- Experience with security best practices for cloud infrastructure.
Benefits
- Competitive salary and equity.
- Comprehensive health, dental, and vision insurance for you and your dependents.
- Opportunity to work on cutting-edge AI technology and make a real impact on the future of software engineering.
- Free lunch and snacks.
Skills
Kubernetes, AWS, GCP, Microvms, Go, Rust, Java, Terraform, Pulumi, Buildkit, CI/CD, Overlaybd
Similar jobs
DevOps / SRE jobsSummer 2027 internship on a Site Reliability Engineering team, building software and automation for deployment, operations, monitoring, and reliability. Requires a software engineering foundation, programming experience, and strong problem-solving and collaboration skills.
Customer-facing DevOps Engineer helping organizations implement secure, compliant cloud infrastructure through the DuploCloud platform. Requires 2–3 years of cloud or DevOps experience, containerization expertise, public cloud knowledge, and strong customer communication skills.
Build and operate large-scale scheduling, storage, caching, and networking infrastructure for AI training and inference. The role targets PhD researchers graduating by December 2026 with systems research depth and strong programming and performance-measurement skills.
Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.
Supports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.