Software Engineer, Fleet Management
Builds systems to manage cloud and bare-metal fleets at scale, integrating hardware metrics with job scheduling and leveraging LLMs for vendor coordination. Requires expertise in cluster and server-level systems like Kubernetes, Terraform, and Linux.
About the job
In this role, you will:
- Design and build systems to manage both cloud and bare-metal fleets at scale.
- Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms.
- Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows.
- Automate infrastructure processes, reducing repetitive toil and improving system reliability.
- Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack.
- Continuously improve tools, automation, processes, and documentation to enhance operational efficiency.
You might thrive in this role if you:
- Have strong software engineering skills with experience in large-scale infrastructure environments.
- Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers).
- Have deep expertise in server-level systems (e.g., systems, containerization, Chef, Linux kernels, firmware management, host routing).
- Are passionate about optimizing the performance and reliability of large compute fleets.
- Thrive in dynamic environments and are eager to solve complex infrastructure challenges.
- Value automation, efficiency, and continuous improvement in everything you build.
Skills
Kubernetes, Terraform, Linux, CI/CD, Chef, Containerization, Cloud Providers
Similar jobs
Fullstack Engineering jobsBuild developer-facing tooling, platform services, and ML infrastructure across Ray and Anyscale, spanning CLI, SDK, APIs, workspaces, observability, and production serving. Requires 5+ years of production software experience, strong systems fundamentals, and familiarity with machine learning tooling.
Build and improve how AI agents discover, interpret, and use Firecrawl by shipping agent-facing product experiences and rigorous A/B tests. The role requires 5+ years of experience, strong product engineering ability, behavioral data fluency, and independent execution.
Own and evolve Firecrawl’s open-source repositories, maintaining the core self-hosting experience and building new context primitives such as a document parsing engine. The role requires 5+ years of experience, a substantial open-source track record, and strong software engineering judgment.
Build and ship developer-facing products that enable AI agents to navigate and act across the web. The role requires 3+ years of experience delivering customer-used products, strong API and developer-experience instincts, and comfort owning ambiguous problems end to end.
Build and own developer-facing document-parsing features that transform complex files into reliable, LLM-ready markdown and structured JSON. The role requires at least three years of shipping products or APIs used by developers, strong product judgment, and comfort solving ambiguous parsing problems.