Skip to content

Senior Runtime Engineer

Design and optimize distributed runtime software for large-scale AI training and inference across heterogeneous clusters. The role requires 3+ years of high-performance or distributed systems experience, strong C/C++ skills, and expertise in concurrency, memory management, and performance optimization.

About the job

Responsibilities

  • Design and implement distributed runtime components to efficiently manage large-scale execution workloads.
  • Develop and optimize high-performance data and communication pipelines that fully utilize CPU, memory, storage, and network resources.
  • Enable scalable execution across multiple compute nodes, ensuring high concurrency and minimal bottlenecks.
  • Collaborate closely with ML and compiler teams to integrate new model architectures, training regimes, and hardware-specific optimizations.
  • Diagnose and resolve complex performance issues across the software stack using profiling and instrumentation tools.
  • Contribute to system design, architecture reviews, and roadmap planning for large-scale AI workloads.

Requirements

  • 3+ years of experience developing high-performance or distributed systems software.
  • Strong programming skills in C/C++, with expertise in multithreading, memory management, and performance optimization.
  • Experience with distributed systems, networking, or inter-process communication.
  • Solid understanding of data structures, concurrency, and system-level resource management, including CPU, I/O, and memory.
  • Proven ability to debug, profile, and optimize code across scales, from threads to clusters.
  • Bachelor's, master's, or equivalent experience in Computer Science, Electrical Engineering, or a related field.

Nice-to-Haves

  • Familiarity with machine learning training or inference pipelines, especially distributed training and large-model scaling.
  • Exposure to Python and PyTorch, particularly for model training or performance tuning.
  • Experience with compiler internals, custom hardware interfaces, or low-level protocol design.
  • Prior work on high-performance clusters, HPC systems, or custom hardware/software co-design.
  • Deep curiosity about unlocking new levels of performance for large-scale AI workloads.

Benefits

  • Work on a breakthrough AI platform beyond the constraints of GPUs.
  • Opportunities to publish and open-source cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Job stability with startup vitality.
  • A non-corporate work culture that respects individual beliefs.
  • Continuous learning, growth, and support.

Skills

C++, C, Python, PyTorch, Distributed Systems, Networking, Inter-Process Communication, Multithreading, Memory Management, Performance Optimization, Data Structures, Concurrency, Profiling, Compiler Internals, Hpc

ZoomInfo

ZoomInfo

Waltham, MA

Senior Software Engineer
$140k+/yrRemote5+ YOEBackend Engineering

Build and scale high-throughput Java backend systems that resolve fragmented data into accurate contact profiles across a 400-million-record platform. The role requires senior software engineering experience with distributed systems, large-scale data processing, and event streaming.

Pindrop

Pindrop

Washington, DC

Senior Software Engineer
$130k+/yrRemote5+ YOEBackend Engineering

Senior Software Engineer owning and evolving high-scale backend and data-layer platform services (IAM, Data Lake, reporting) that power all Pindrop products. Design, build, and operate distributed systems on AWS/GCP with strong focus on reliability, scalability, on-call, and mentorship.

Mintlify

Mintlify

San Francisco, CA

Senior Backend Engineer
$190k+/yrOn-site5+ YOEBackend Engineering

Build and scale backend APIs, content pipelines, and AI infrastructure powering Mintlify’s documentation products. The role requires 4+ years of software development experience, strong systems and production engineering skills, product judgment, and high ownership.

Cape

Cape

New York, NY

Lead Distributed Systems Engineer
$250k+/yrHybrid7+ YOEBackend Engineering

Lead the architecture, implementation, and engineering team responsible for reliable, secure data synchronization across cloud and intermittently connected private 5G edge networks. The role requires deep distributed-systems expertise, event-driven architecture experience, and demonstrated technical leadership.

Okta

Okta

Toronto, Canada

Senior Software Engineer, Device Identity
CA$136k+/yrHybrid5+ YOEBackend Engineering

Senior software engineer designing and scaling secure server-side services for Okta’s device identity platform. The role requires strong Java and Spring expertise, database and API design experience, and the ability to mentor engineers while driving scalable, secure development practices.