Design and build core infrastructure for AI inference on Cloudflare's global network of GPUs and accelerators. Optimize scheduling, routing, reliability, and observability for low-latency, serverless LLM and model serving at the edge. Requires expert Rust and distributed systems experience.
Salary not listed
On-site5+ YOEML Engineering
About the role
Responsibilities
Develop and maintain core components of the serverless inference platform to ensure high availability and scalability.
Optimize the model scheduling system to increase efficiency and resource utilization across inference infrastructure.
Implement improvements to request routing logic to reduce latency for end-users.
Drive measurable improvements in platform reliability and resilience by identifying and mitigating systemic risks.
Expand and refine the observability stack (metrics, logging, tracing) and fine-tune alerts to proactively identify and resolve production issues.
Lead complex, cross-functional technical projects from concept and design through deployment and operationalization.
Mentor junior engineers and contribute to a strong, collaborative engineering culture.
Requirements
Proven experience in systems engineering with a primary focus on distributed, high-performance systems.
Expert proficiency in Rust programming, particularly in an asynchronous environment.
Deep understanding and hands-on experience with networking and application protocols (e.g., TCP, HTTP, WebSocket).
Solid experience with scaling and performance optimization techniques, including load balancing and caching in a distributed environment.
Nice-to-Haves
Demonstrable experience with container orchestration platforms, specifically Kubernetes and/or Nomad.
Familiarity with the unique architectural challenges involved in large-scale inference serving (e.g., LLMs and diffusion models).
Senior Applied ML Engineer building and shipping user-facing LLM and generative AI features. Own architecture for production ML systems, partner cross-functionally with product/research teams, and establish best practices for prompt engineering and model evaluation. Requires 4+ years software engineering with 2+ years on ML products, strong backend skills.
159k – 232k/yr
Remote4+ YOEML Engineering
Sr. Software Engineer
DialpadSan Francisco, CA
Lead the architecture and development of Dialpad's autonomous Agentic AI platform, building multi-agent orchestration, memory systems, real-time reasoning, and tool execution for enterprise workflows. Requires 10+ years experience, prior technical leadership at Staff/Principal level, and deep expertise in LLM platforms, agent frameworks, and production AI infrastructure.
225k – 252k/yr
On-site10+ YOEML Engineering
Lead Research Engineer, Data Quality
hudSan Francisco, CA
Lead the data quality team at HUDHUD to build QC systems, validation methods, and experiments that measure and improve training data for frontier AI agents and RL environments. Requires deep data quality intuition, Python/Docker/Linux proficiency, and experience turning research insights into production evaluation pipelines.
Salary not listed
On-site7+ YOEML Engineering
Senior Software Engineer, Perception
Shield AIWashington, DC +3
Develop and deploy advanced vision, VLM, and VLA machine learning models for autonomous systems perception. Own models from training through optimization and deployment on embedded hardware, collaborating with research and engineering teams to deliver production capabilities for complex real-world environments.
163k – 245k/yr
On-site5+ YOEML Engineering
Lead Data Scientist
Apartment ListUnited States
Lead Data Scientist building and deploying production ML models for a two-sided rental marketplace. Own end-to-end projects in ranking, personalization, renter intent, demand and supply optimization using Python, SQL, and ML frameworks. 4+ years experience required.