Staff Software Engineer, AI Data Platform
Build and scale eval systems, fine-tuning pipelines, agent-first infrastructure, and data platforms that power frontier AI labs at Labelbox. Requires full-stack prototyping expertise, strong architecture judgment, daily use of coding agents, and deep TypeScript/Python proficiency (7+ years implied by Staff level).
About the job
What you may work on
- Eval systems that run millions of agent trajectories to measure model and product quality.
- Fine-tuning pipelines that turn evaluation signals into measurable agent improvements.
- Agent-first product surfaces: UX and infrastructure for workflows where the user is a model or an agent operator.
- The systems behind hundreds of thousands of AI interviews used to source and match freelance workers to projects.
- Infrastructure that scales to the throughput frontier labs actually need.
- Integration of the latest models and capabilities into production within days of release.
What we're looking for
- 4+ year track record of shipping systems customers and other engineers rely on
- You build full stack prototypes fast and they hold up. The v1 you ship becomes the foundation the rest of the team builds on.
- Strong system and API design judgement
- Hard architecture and product calls land with you. You make them, defend them under pressure, and update fast when someone else is right.
- You ship production code with coding agents daily. You know where they break and what it takes to make them reliable to further accelerate the team's velocity.
- You set direction by being the example. Other engineers reach for your designs and your code as the reference.
- You move fast in ambiguous, startup-pace environments with influence over authority.
- You have worked in all parts of the stack
- Deep proficiency in TypeScript and/or Python.
Nice to have
- Production experience building LLM- or agent-driven products.
- Designing evaluations for LLMs and agents, or producing high-quality data for ML systems.
- Background in production distributed systems, ML infrastructure, or data systems at scale.
Our Technology Stack
- Frontend: React.js with Redux, TypeScript
- Backend: Node.js, TypeScript, Python, some Java & Kotlin
- APIs: GraphQL
- Cloud & Infrastructure: Google Cloud Platform (GCP), Kubernetes
- Databases: MySQL, Spanner, PostgreSQL
- Queueing / Streaming: Kafka, PubSub
Compensation
Annual base salary range: $250,000–$280,000 USD (not inclusive of equity or benefits; varies based on skills, experience, and location).
Skills
TypeScript, Python, React, Node.js, GraphQL, Kubernetes, GCP, Postgres, MySQL, Kafka, LLMs, Distributed Systems
Similar jobs
ML Engineering jobsStaff Machine Learning Engineer building and operating production ML systems for causal marketing measurement, optimization, and planning. The role requires deep statistical and machine learning expertise, production programming experience, cross-functional collaboration, and technical mentorship.
Staff engineer responsible for designing and scaling the infrastructure, execution environments, verifiers, and tooling used to train and evaluate AI agents. Requires 8+ years of software engineering experience, strong Python and distributed-systems expertise, and familiarity with sandboxing, high-throughput systems, and LLM workflows.
Senior Staff ML Engineer fine-tunes and optimizes state-of-the-art LLMs for Airbnb's customer support AI products, including AI assistants and autonomous agents. Partners cross-functionally to productionize models at scale. Requires PhD and 10+ years experience with PyTorch.
Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.
Leads the roadmap and technical vision for Snowflake Feature Store, building reliable, high-performance machine learning platform capabilities and supporting technical execution across partner teams. Requires 10+ years of experience with data-serving infrastructure or ML platforms, plus Java and Python expertise.