Software Engineer
Build and operate realtime and batch data pipelines processing billions of events daily at xAI. Design distributed data platforms, own data correctness, create shared datasets for product and business teams, and partner on data acquisition using tools like Spark, Kafka, Flink, and SQL.
About the job
Responsibilities
- Design, build, and operate production-grade realtime and batch pipelines that ingest, process, validate, and deliver data powering user-behavior insights and product decisions.
- Create shared datasets, fact tables, and internal data products that let other teams analyze, debug, and improve product performance.
- Prototype and build tooling that automates and accelerates internal data workflows — backfills, dashboards, report generation, and self-serve access to data.
- Own data correctness end to end: validate with output invariants, denominator reconciliation, and independent recomputation, and lead root-cause investigations when key metrics move unexpectedly.
- Move fluidly across query engines and frameworks (e.g., BigQuery, Trino, Clickhouse for analytics; Flink, Kafka, Spark/Scalding for streaming and batch), choosing the right tool and adapting quickly to new infrastructure and environments.
- Partner across product and business teams to surface where data gaps exist and prioritize the highest-impact opportunities for new data acquisition and improvement.
- Iterate quickly on feedback, shipping the smallest useful increment with a strong bias toward efficient, accurate, and reliable solutions.
Requirements
- 3+ years of professional software engineering experience, ideally in data engineering or distributed systems.
- Hands-on expertise in Python, Rust, Scala, Go or Java, and data pipeline toolings and distributed systems.
- Knowledge of realtime and batch data processing tools such as Spark/Kafka/Flink/SQL and various storage systems in RMDBs/NoSQL.
- Experience solving large scale problems and comfortable doing incremental quality work while building brand new systems to enable future quality improvements.
- Proven records of interpreting product requirements into engineering implementation plans, and effectively communicating with different groups (AI, product, marketing/sales and engineering).
Compensation and Benefits
- $125,000 - $400,000 USD total compensation (base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks).
Skills
Python, Rust, Scala, Go, Java, Spark, Kafka, Flink, SQL, BigQuery, Trino, ClickHouse, Distributed Systems, Data Pipelines
Similar jobs
Data Engineering jobsBuild and scale AWS-based data infrastructure, pipelines, knowledge graphs, and APIs for a CTV performance advertising platform. The role requires production data engineering experience with Spark, Scala, AWS, SQL, and large-scale services, plus a bachelor's degree.
Build and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.
Build scalable analytics engineering infrastructure, SaaS data models, and AI-enabled workflows that support enterprise decision-making. The role requires 3–6 years of hands-on analytics or data engineering experience, strong SQL and modern data modeling expertise, and cloud data warehouse experience.
Build and operate the data platform supporting automated regulatory reporting for a prediction markets business. The role combines SQL and dbt development, end-to-end data investigation, automated validation, and cross-functional ownership under strict deadlines.