Senior Software Engineer, Infrastructure & Platform
Designs and builds scalable core infrastructure for data generation, human-in-the-loop workflows, and AI evaluation pipelines. Requires strong experience in distributed systems, cloud platforms like GCP/AWS, and high-throughput data processing.
About the job
Responsibilities
- Design and build core infrastructure systems.
- Architect and develop the shared infrastructure powering our data generation platforms, human-in-the-loop systems, and evaluation pipelines.
- Build systems capable of processing large-scale datasets and high-throughput workloads with strong reliability guarantees.
- Create reusable infrastructure and APIs that enable product engineers and researchers to build quickly and reliably on top of core systems.
- Design systems with strong observability, monitoring, and fault tolerance to support production workloads at scale.
- Help define long-term system architecture across data pipelines, compute infrastructure, task orchestration, and storage systems.
- Work closely with engineers and researchers to support new AI experimentation workflows and platform capabilities.
- Define standards for system design, deployment, reliability, and infrastructure operations.
Required Qualifications
- Strong experience building production distributed systems or platform infrastructure.
- Proficiency in Python and/or JavaScript (Node.js / Next.js) or similar backend technologies.
- Experience designing and operating systems in cloud environments (GCP or AWS).
- Experience with message queues and event-driven systems (Kafka, RabbitMQ, Pub/Sub, etc.).
- Experience working with high-throughput data pipelines and asynchronous processing systems.
- Strong understanding of system scalability, performance, and reliability.
- Experience owning systems running in production environments.
Preferred Qualifications
- Experience building internal developer platforms or shared infrastructure.
- Experience supporting large-scale data processing pipelines.
- Experience with AI infrastructure, LLM evaluation systems, or ML pipelines.
- Experience working at high-growth startups or scaling early infrastructure.
- Experience designing human-in-the-loop or workflow orchestration systems.
Skills
Python, JavaScript, Node.js, Next.js, GCP, AWS, Kafka, RabbitMQ, Pub/Sub, Distributed Systems
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.