Build and operate reliable infrastructure and testing services for autonomous-vehicle development. The role requires 5+ years supporting production services and SRE responsibilities, plus proficiency in Python or Golang and experience with automation, observability, CI/CD, and resilient infrastructure.
165k – 208k/yr
Hybrid5+ YOEDevOps / SRE
About the role
Responsibilities
Measure and maintain the uptime of services critical to autonomous-vehicle development, including testing and validation of on-vehicle software for hardware platforms.
Support all phases of service delivery, including design, deployment, operations, support, automation, and continuous improvement.
Ensure the uptime, stability, accuracy, and usability of robot testing platforms.
Collaborate with other teams to support adoption of testing frameworks and maintain key tests.
Work with systems handling large data volumes, data-processing pipelines, and compute-intensive CPU and GPU workloads.
Requirements
Bachelor's degree in Engineering, Computer Science, Mathematics, or a related field.
5+ years supporting production services, participating in on-call rotations, and performing SRE responsibilities.
Proficiency in Python or Golang.
Experience building and managing infrastructure services.
Experience writing system-level automation, building resilient infrastructure, and developing CI/CD pipelines.
Experience building observability stacks, instrumentation, OpenTelemetry, and Grafana dashboards.
Experience building software services, writing APIs for backend services, and owning and managing full-stack applications.
Nice-to-Haves
Linux system administration experience, including kernel troubleshooting and device-driver development.
Continuous deployment and release-management experience.
Experience with CI toolchains such as Buildkite and Bazel.
Experience with test frameworks such as Pytest.
Experience writing and managing infrastructure with infrastructure-as-code tools such as Terraform, Ansible, and Salt.
Skills
PythonGoLinuxCI/CDOpenTelemetryGrafanabuildkiteBazelpytestTerraformAnsiblesaltAPIsInfrastructure As Codedevice drivers
Infrastructure Engineer responsible for securing, scaling, and maintaining cloud architecture, Kubernetes clusters, ML pipelines, and CI/CD systems at a computer vision AI startup. Must have production Kubernetes, IaC, and AWS/GCP experience.
165k – 200k/yrHybridDevOps / SRE
DevOps Engineer
OctusNew York, NY
Design, implement, and maintain cloud infrastructure and CI/CD pipelines. Collaborate with developers, SRE, and Security to ensure system reliability, scalability, and security.
165k – 190k/yrOn-site5+ YOEDevOps / SRE
OS / K8s Systems Engineer
BasetenSan Francisco, CA +1
Build automation and systems to provision and orchestrate GPU hardware into scalable Kubernetes clusters. Requires deep Linux expertise, provisioning experience, and strong programming in Python/Go.
165k – 330k/yrHybridDevOps / SRE
Site Reliability Engineer (SRE)
BasetenSan Francisco, CA +1
Site Reliability Engineer builds and maintains scalable infrastructure for ML model deployment, automates CI/CD pipelines, and ensures reliability using tools like Kubernetes and Terraform. Collaborates cross-functionally, owns projects end-to-end, and mentors juniors; bachelor's in CS or related field required.
165k – 330k/yrHybridDevOps / SRE
Software Engineer - Internal Platform
BasetenSan Francisco, CA +1
Builds internal tooling, monorepos, CI/CD pipelines, and shared libraries to boost engineering productivity. Requires strong proficiency in Go/Python, Kubernetes/Docker experience, and monorepo management.