Senior engineer building and improving Bazel-based build, test, and packaging tools for Datadog's large multi-language monorepo. Own projects end-to-end to boost developer productivity, performance, and CI efficiency at massive scale.
244k – 305k/yr
Hybrid5+ YOEDevOps / SRE
About the role
What You’ll Do
Invent build, test and packaging tools that are simpler and more reliable to use.
Push performance and cost efficiency at scale, raising cache hit rates and cutting CI times and compute spend across millions of targets.
Treat CI like SREs treat prod, making sure our pipelines are green and fast.
Prepare, run, and finish complex migrations.
Contribute back to the Bazel ecosystem, upstreaming fixes and shaping features we depend on.
Who You Are
An expert in Bazel and/or one of the languages listed above.
A well-rounded engineer. You must broadly understand the various types of software projects that are built, tested, and packaged with our tools.
Both careful and fearless. The changes we make impact the velocity of hundreds of engineers. They are risky but necessary.
User-focused. We help Datadog engineers to use the tools that we develop, and continuously improve their usability, so they don’t need our help the next time.
Ideally, you have experience working with a large codebase.
Lead the network production engineering team responsible for availability, performance, and automation of fabrics supporting 100k+ accelerator clusters at massive scale. Own SLOs, build remediation automation, and set operating models between design and site teams.
242k – 284k/yr
On-site7+ YOEDevOps / SRE
Senior Software Engineer, Infra/Systems
ConvexSan Francisco, CA
Convex is seeking a Senior Software Engineer to design, build, and maintain their global cloud infrastructure. This role involves working on core systems, improving performance and reliability, and owning architectural decisions.
240k+/yr
Hybrid6+ YOEDevOps / SRE
Software Engineer, Frontier Systems
OpenAISan Francisco, CA
Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.
250k – 445k/yr
On-site7+ YOEDevOps / SRE
Data Center Controls Network Engineer
OpenAISan Francisco, CA
Designs, validates, and scales secure OT network architectures for high-density AI data centers, including controls systems, telemetry, and integration with IT infrastructure. Requires 8+ years in OT networking, industrial protocols, and resilient topologies in mission-critical environments.
257k – 327k/yr
Hybrid8+ YOEDevOps / SRE
Site Reliability Engineer
Forward NetworksSanta Clara, CA
Build the Site Reliability Engineering function from the ground up at Forward, defining SLOs, building observability infrastructure, leading incident response, and embedding reliability into the SDLC for their complex SaaS platform. Requires 6+ years SRE/DevOps experience, strong networking and Kubernetes skills, and a track record maturing SRE practices.