Skip to content
AnyscaleAnyscale

Staff Software Engineer, Ray Data

Staff Software Engineer responsible for designing and scaling Ray Data’s distributed data-processing infrastructure for large-scale AI training and inference. Requires 6+ years of production software and architectural ownership experience, plus deep distributed-systems expertise and strong Python skills.

About the job

Responsibilities

  • Design, build, and improve the core systems powering Ray Data, focusing on performance, scalability, and reliability.
  • Design and optimize distributed execution across data-pipeline stages in heterogeneous environments.
  • Build data-loading and processing solutions for production training and inference workloads.
  • Solve problems involving distributed execution, scheduling, resource management, data partitioning, fault tolerance, and performance optimization.
  • Make system-level architectural decisions across resource allocation, execution models, batch versus streaming workloads, and consistency versus availability.
  • Work with customers and AI-native companies to address challenges in scaling AI workloads.

Requirements

  • 6+ years of experience building production-grade software, infrastructure, or developer-facing systems.
  • Strong Python engineering experience.
  • 6+ years personally owning core architectural decisions within a distributed data or compute engine.
  • Deep experience with distributed-systems internals, including scheduling, fault tolerance, data partitioning, distributed execution, performance optimization, or database and query-engine internals.
  • Demonstrated ability to reason through system-level tradeoffs and defend architectural decisions.
  • Passion for large-scale AI infrastructure.

Compensation and Benefits

  • $240,000–$270,000 annual salary.
  • Equity and health, dental, and vision coverage, with many plans up to 99% employer-covered.
  • Flexible time off, paid parental leave, and mental health support.

Skills

Python, Ray Data, Distributed Systems, Distributed Execution, Scheduling, Resource Management, Data Partitioning, Fault Tolerance, Performance Optimization, Database Internals, Query Engines, Batch Processing, Streaming Processing, AI Infrastructure

Haus

Haus

San Francisco, CA
Staff Backend Engineer - Data Platform
$240k+/yrHybrid10+ YOEData Engineering

Staff-level engineer leading backend services and data-platform architecture, including large-scale ingestion, distributed systems, and trustworthy BigQuery/dbt warehouse models. Requires 10+ years of software engineering experience, expert Python, deep SQL/dbt expertise, and strong technical leadership.

Snowflake

Snowflake

Menlo Park, CA

Staff Software Engineer - Snowhouse
$236k+/yrOn-site12+ YOEData Engineering

Leads the design and operation of highly available distributed data platforms and pipelines at Snowflake, while providing technical leadership across teams. Requires 12+ years of distributed-systems experience, cloud expertise, and strong database and system-design depth.

Harvey

Harvey

San Francisco, CA
Staff Software Engineer, Data Platform
$231k+/yrHybrid10+ YOEData Engineering

Staff Software Engineer on the central data platform team, responsible for architecture, ingestion, orchestration, streaming, governance, and self-service data tooling. Requires 10+ years building production data infrastructure and deep experience with warehouses, CDC, streaming, orchestration, Python, and SQL.

Vercel

Vercel

San Francisco, CA
Staff Data Platform Engineer - Finance
$260k+/yrHybrid8+ YOEData Engineering

Staff Data Platform Engineer leading the architecture and development of financial data infrastructure for revenue reporting, billing, forecasting, and compliance. Requires 8+ years of data engineering or architecture experience, strong streaming and warehouse expertise, and the ability to mentor engineers and partner with Finance and Audit leaders.

OpenAI

OpenAI

San Francisco, CA
Analytics Engineer, GTM
$220k+/yrHybrid10+ YOEData Engineering

Analytics Engineer supporting Go-to-Market teams by building scalable data models, metrics, pipelines, visualizations, and self-service products. The role requires 10+ years of data experience, deep SQL expertise, Python proficiency, and strong business judgment.