Software Engineer, Dev Velocity
Build internal developer platform, tooling, and automation to accelerate engineering velocity. Focus on CI/CD pipelines, test infrastructure, build systems, and metrics to help engineers ship faster and more reliably.
About the job
What You'll Do
- Improve speed and reliability of early change feedback
- Invest in build and dependency tooling, caching, incremental builds, and monorepo ergonomics
- Accelerate CI/CD pipelines - cutting build, test, and deploy times while keeping the path to production safe and reversible
- Make test suites fast and trustworthy: parallelization, smart test selection, flaky-test detection, and better local/preview environments
- Design and build tools to support concurrent agent execution
- Make test feedback fast enough to live inside an agent's loop
- Get agents the right context (codebase conventions, service ownership, recent changes, environment state)
- Simplify and establish repeatable patterns
- Reduce toil through automation, paving over repetitive workflows and manual steps
- Instrument the engineering org with meaningful signals (deployment frequency, lead time for changes, change failure rate, time to restore)
- Build and participate in an on-call rotation
What We're Looking For
- Experience building and operating developer-facing tooling, CI/CD systems, or internal platforms that other engineers depend on every day
- Experience driving cultural and process changes to make engineering teams more effective
- Strong software engineering fundamentals and a track record of shipping, maintaining, and debugging production systems
- Proficiency in Go (or a strong willingness to ramp quickly)
- Hands-on experience with Kubernetes or similar container orchestration
- Familiarity with infrastructure-as-code tools (Terraform)
- A bias toward measurement: instrument what you build and let data tell you whether you actually made things faster
- Empathy with engineers, fixing developer pain points in an elegant and sustainable manner
Nice-to-Haves
- Experience improving build systems, test infrastructure, or developer experience at meaningful scale (Bazel, Concourse, GitHub workflows)
- Familiarity with observability tools like Datadog, Grafana, and OpenTelemetry
- Comfort working close to the metal - Linux, containers, kubernetes
- Security-minded instincts, especially around supply chain and the path from commit to deploy
- Understanding of where to use AI as an accelerant, and where human context must come first
Benefits
- 4 weeks of paid vacation
- 14 weeks of fully paid parental leave
- Long-term disability, life insurance, and 401K plans
- 100% employer-paid medical coverage and 99% employer-paid dental and vision coverage
- FSAs and HSAs available
- Monthly lifestyle stipend for wellness, mental health and therapy, hobbies
- Monthly cell phone and internet subsidy
- Commuter benefits for Bay Area, home office stipends for remote
- Continuous learning benefits & related support
Skills
Go, Kubernetes, Terraform, CI/CD, Bazel, Concourse, GitHub Actions, Datadog, Grafana, OpenTelemetry
Similar jobs
DevOps / SRE jobsLeads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.