
Anyscale
San Francisco, CA
Scalable compute platform for AI built on Ray
About
Anyscale provides a fully-managed platform based on the open-source Ray framework for developers to build, scale, and run distributed AI and ML applications from laptops to cloud clusters. It serves AI teams at companies like OpenAI, Uber, and Amazon, eliminating infrastructure management to focus on innovation. This enables effortless scaling of AI workloads critical for production AI.
Tech stack
Kubernetes, AWS, GCP, Azure, Python, Terraform, GitHub Actions, PyTorch, Go
Perks & benefits
Health insurance, Dental insurance, Vision insurance, Fertility benefits, Parental leave, Unlimited PTO, Learning stipend, Wellness stipend, Free lunch, Mental health support
More AI companies
AI companiesSan Francisco, CA
San Francisco, CA
San Francisco, CA
Redwood City, CA
Cambridge, MA
San Francisco, CA
Open jobs
20Staff Software Engineer responsible for designing and scaling Ray Data’s distributed data-processing infrastructure for large-scale AI training and inference. Requires 6+ years of production software and architectural ownership experience, plus deep distributed-systems expertise and strong Python skills.
Develop and improve Ray Core’s C++ distributed-systems backend, focusing on performance, reliability, fault tolerance, and scalability. The role requires at least five years of experience with distributed systems, C/C++, low-level operating systems, algorithms, and system design.
Build developer-facing tooling, platform services, and ML infrastructure across Ray and Anyscale, spanning CLI, SDK, APIs, workspaces, observability, and production serving. Requires 5+ years of production software experience, strong systems fundamentals, and familiarity with machine learning tooling.
Own detection engineering and lead incident response across corporate and production environments, building cloud, endpoint, runtime, and Kubernetes coverage. The role requires 6+ years in security, hands-on detection development, and end-to-end incident leadership.
Own Anyscale’s secure software development lifecycle, partner with engineering on secure architecture and features, and lead vulnerability management and remediation. The role requires 8+ years of product or application security experience and strong hands-on secure-development expertise.
Own and advance the security of Anyscale’s production and multi-cloud infrastructure, including hardening, segmentation, Kubernetes runtime protection, and access controls. Requires 8+ years of security engineering experience and hands-on expertise with AWS, Azure, Kubernetes, and cloud security tooling.
Own Anyscale’s compliance function end to end, leading SOC 2 and ISO 27001 programs, audit readiness, customer security diligence, and enterprise risk management. The role requires 7+ years in governance, risk, and compliance plus strong cloud and SaaS security-controls expertise.
Own Anyscale’s hands-on IT function across identity, endpoint management, hardware logistics, SaaS vendors, audit controls, and internal automation. The hybrid role requires strong Okta and Google Workspace administration, macOS fleet management, HRIS-driven identity automation, and an automate-first mindset.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.
Builds and scales control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, Kubernetes deployments, and accelerator integration. Requires 3+ years of production software experience, distributed-systems expertise, and proficiency in Go and Python.
Build and optimize Ray’s distributed data-processing and Datasets libraries, working across Ray Core, Apache Arrow, machine-learning integrations, and streaming workloads. The role requires at least five years of relevant experience and strong distributed-systems, algorithms, and data-processing expertise.
Build and improve Ray’s C++ distributed-systems backend, focusing on performance, reliability, fault tolerance, testing, and scalable programming. The role requires at least two years of relevant experience plus strong algorithms, data structures, system design, and distributed-systems expertise.
Lead Infrastructure, SRE, and Enterprise Governance teams at Anyscale to build and scale distributed computing platform for Ray, including cluster management, Kubernetes support, autoscaling, reliability, and cloud integrations. Requires strong engineering management experience, deep distributed systems knowledge, and Kubernetes expertise.
Software Engineer building scalable control and data plane infrastructure for Anyscale's Ray platform. Design and optimize cluster orchestration, scheduling, Kubernetes deployments, and accelerator support for distributed AI/ML workloads. Requires 3+ years production experience with cloud-native tech, Go/Python, and distributed systems.
Enterprise Account Executive responsible for prospecting, developing, and closing new business for Anyscale's Ray-based ML platform. Requires 5+ years full-cycle software/cloud sales experience with emphasis on ML, cloud, and infrastructure solutions.
Lead Anyscale's Customer Engineering team to deliver technical support for production AI workloads on their Ray-based platform. Partner with Product and Engineering to drive improvements from customer feedback while scaling operational excellence, automation, and self-service.
The Customer Engineer will help customers onboard, adopt, and grow on the Anyscale platform, troubleshooting and resolving technical issues. This role involves close coordination with engineering teams to debug complex issues and contribute to product improvement.
Build and optimize distributed LLM inference systems at scale using Ray, integrating with engines like vLLM to deliver high-throughput, low-latency batch and online inference solutions.
Leads product roadmap for Ray Data, balancing open source growth with commercial features for Anyscale Runtime. Requires 4+ years PM experience, strong technical background in distributed systems/ML infrastructure, and strategic thinking for ML/data lifecycle.