Full Stack Engineer, Fleet Scheduling
Develops full-stack web applications for real-time monitoring, job scheduling, and resource management of AI supercomputing workloads. Requires expertise in modern frontend/backend technologies, APIs, distributed systems, and cloud infrastructure.
About the job
Responsibilities
- Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time.
- Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends.
- Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation.
- Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability.
- Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems.
- Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure.
Requirements
- Significant experience in full-stack development, with expertise in modern frontend frameworks (React, Vue, or Angular) and backend technologies (Python, Go, or Node.js).
- Experienced in building scalable, high-performance web applications for complex distributed systems.
- Strong understanding of RESTful and GraphQL APIs, distributed databases, and cloud infrastructure (especially Azure).
- Execution-focused with a keen eye for usability, performance, and scalability in enterprise-scale systems.
- Comfortable working in fast-paced, highly collaborative environments with tight timelines and evolving priorities.
Nice-to-Haves
- Experience working with Kubernetes, Docker, and cloud-native application deployment.
- Understand AI/ML workload scheduling and orchestration challenges.
- Experience with real-time data processing, visualization libraries, and observability tooling.
Skills
React, Vue, Angular, Python, Go, Node.js, REST APIs, GraphQL, Kubernetes, Docker, Azure
Similar jobs
Fullstack Engineering jobsBuild developer-facing tooling, platform services, and ML infrastructure across Ray and Anyscale, spanning CLI, SDK, APIs, workspaces, observability, and production serving. Requires 5+ years of production software experience, strong systems fundamentals, and familiarity with machine learning tooling.
Build and improve how AI agents discover, interpret, and use Firecrawl by shipping agent-facing product experiences and rigorous A/B tests. The role requires 5+ years of experience, strong product engineering ability, behavioral data fluency, and independent execution.
Own and evolve Firecrawl’s open-source repositories, maintaining the core self-hosting experience and building new context primitives such as a document parsing engine. The role requires 5+ years of experience, a substantial open-source track record, and strong software engineering judgment.
Build and ship developer-facing products that enable AI agents to navigate and act across the web. The role requires 3+ years of experience delivering customer-used products, strong API and developer-experience instincts, and comfort owning ambiguous problems end to end.
Build and own developer-facing document-parsing features that transform complex files into reliable, LLM-ready markdown and structured JSON. The role requires at least three years of shipping products or APIs used by developers, strong product judgment, and comfort solving ambiguous parsing problems.