Software Engineer, Habitat (Online Data)
Builds and operates Habitat, OpenAI's core online database platform handling high-QPS, latency-sensitive workloads. Owns end-to-end distributed systems for storage, caching, routing, CDC, and privacy; requires 8+ years experience with Rust/Python expertise.
About the job
Responsibilities
- Design and build core abstractions spanning storage, caching, routing, CDC, and privacy enforcement
- Own a major surface area end to end, from product and API design to operational excellence
- Improve latency, correctness, and cost efficiency for real production workloads at massive scale
- Build strong instrumentation, debugging workflows, and developer-first tooling
- Collaborate closely with internal product and infrastructure teams to understand requirements and ship pragmatic solutions
- Participate in an on-call rotation and raise the bar on reliability while aggressively improving performance and usability
Requirements
- Strong track record building and operating high-scale backend or data-intensive distributed systems in production
- Excellent systems judgment and the ability to make tradeoffs across latency, cost, correctness, and reliability
- Deep care for developer experience: simple abstractions, strong defaults, guardrails, and debuggability
- Comfort owning ambiguous problems end to end and driving roadmap plus execution
- Experience in one or more of: databases, caching systems, routing and load balancing, indexing and retrieval, CDC pipelines
- Deep experience with tail latency and global performance optimization (p95, p99, request steering, locality)
- Experience designing platform APIs consumed by many internal teams
- Experience with multi-region systems, consistency semantics, and failover design
- 8+ years of industry experience building production software, including 3+ years leading large-scale, complex projects or technical initiatives as a tech lead or senior IC
- Strong passion for building distributed systems at scale, with a focus on reliability, scalability, security, and continuous improvement
- Proficiency in Rust and/or Python (Rust preferred for core systems work; Python commonly used for tooling, services, and ecosystem integration)
- Excellent communication skills, with the ability to build alignment and drive decisions across diverse technical and non-technical stakeholders
Skills
Rust, Python, Distributed Systems, Databases, Caching, Routing, Cdc, Privacy Enforcement, APIs, Observability
Similar jobs
Data Engineering jobsLeads the architecture, scaling, security, governance, and cost optimization of enterprise and AI data platforms. Requires 10+ years of data or software engineering experience, with expertise in production data foundations, CI/CD, governance, security, and performance optimization.
Build and scale data pipelines, reusable datasets, and validation frameworks supporting business intelligence, marketing, and data science. The role requires strong Python and SQL skills, modern data-stack experience, and at least four years of software or data engineering experience.
Own the reliability, performance, observability, scalability, and cost efficiency of large Aurora MySQL production environments supporting healthcare applications. The role requires 6+ years of database engineering experience, deep MySQL and AWS expertise, and strong skills in automation, incident response, and query optimization.
Leads an analytics engineering team that transforms raw data into reliable, actionable insights for product, marketing, and operations. The role requires 7+ years in data or analytics engineering, management experience, and advanced SQL, Databricks, and dbt expertise.
Senior Data Infrastructure Engineer responsible for building and operating reliable, low-latency streaming and batch data systems that support AI products. Requires 5+ years of production data infrastructure experience and expertise with technologies such as Kafka, Flink, ClickHouse, and Terraform.