Software Engineer, Infrastructure - Analytics Platform
Owns end-to-end production-critical infrastructure for analytics platform, building performant backend systems in Rust or C++ and operating distributed services at scale on Kubernetes. Requires strong systems experience in performance optimization, debugging, and on-call reliability.
About the job
Responsibilities
- Own critical infrastructure across design, implementation, rollout, operation, and iteration.
- Build and operate performant backend systems in Rust or C++ that support core research workflows.
- Design and improve distributed data and serving systems, including tradeoffs around partitioning, replication, consistency, retries, backpressure, and failure isolation.
- Debug real production bottlenecks across latency, throughput, contention, hot spots, and overload behavior.
- Operate business-critical services through on-call, incidents, postmortems, observability, rollout safety, and zero-downtime migrations.
- Improve reliability of services running on Kubernetes, including resource tuning and failure handling.
- Partner closely with engineers and researchers to deliver fast, reliable, useful systems.
- Raise the bar through strong technical judgment, ownership, and follow-through.
Requirements
- Track record of owning operationally critical systems end to end and delivering outcomes in ambiguous environments.
- Strong hands-on experience building performance-sensitive backend systems in Rust or C++.
- Comfort working below typical service abstractions, including concurrency, async execution, memory behavior, serialization, I/O, networking, profiling, and failure analysis.
- Experience designing, building, or operating distributed systems or distributed databases at meaningful scale.
- Hands-on experience operating production-critical systems, including incidents, observability, rollout safety, and recurrence prevention.
- Strong judgment in balancing engineering quality, speed, risk, and business impact.
- Habit of shipping practical first versions and improving them through production feedback.
Preferred
- Experience with ClickHouse-like systems or infrastructure for analytics, telemetry, logging, search, ingestion, storage, or query execution.
Skills
Rust, C++, Kubernetes, Distributed Systems, ClickHouse, Observability, Profiling, Networking, Concurrency, Async Execution
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.