Latest DevOps / SRE jobs
Job results
Design, build, and operate highly available distributed cloud infrastructure for a multi-cloud SaaS platform. The role requires 5+ years of software development experience, cloud and infrastructure-as-code expertise, and strong knowledge of networking and security.
Senior software engineer responsible for designing, automating, and operating secure, highly available distributed cloud infrastructure across multi-cloud environments. Requires 5+ years of experience with scalable systems, cloud platforms, infrastructure as code, and production debugging.
Design and operate highly available, distributed cloud infrastructure for a multi-cloud ClickHouse Cloud platform. The role requires 5+ years of software development experience, strong systems and cloud expertise, and proficiency with infrastructure automation and security.
Senior software engineer designing, deploying, and operating scalable cloud infrastructure across distributed, multi-cloud environments. Requires 5+ years of experience with distributed systems, cloud platforms, infrastructure as code, Kubernetes, networking, and cloud security.
Design, build, and operate highly available distributed cloud infrastructure for a multi-cloud ClickHouse Cloud platform. The role requires 5+ years of software development experience, cloud infrastructure expertise, and strong knowledge of networking and security.
Infrastructure Engineer scales ML inference systems serving 150+ biological models using Kubernetes and AWS. Requires containerization expertise, cloud knowledge, and onsite presence in San Francisco.
DevOps Engineer builds and maintains CI/CD pipelines, ML model infrastructure, and automated testing for AI image/video software products. Requires 2+ years experience, C++ build tools expertise, and cloud platforms like AWS/Azure.
Develops internal tools and infrastructure for autonomy lifecycle testing, replay systems, and diagnostics in robotics. Requires 3+ years experience with C++, Python, simulation frameworks, and performance-sensitive systems.
Builds and operates reliable, scalable AI infrastructure including observability, SLOs, incident response, automation, and performance tuning for ultra-low-latency serverless compute. Requires 3+ years SRE/DevOps experience with cloud, Kubernetes, programming (Go/Rust/Python), and observability tools.
Senior Infrastructure Engineer owns critical infrastructure decisions, builds scalable platforms using AWS and Terraform, ensures security/compliance, and mentors teams. Requires 5+ years AWS experience and expertise in monitoring tools like Datadog.
Site Reliability Engineer who owns production services end-to-end, writes production code, builds observability and internal tooling, and embeds with product teams to improve reliability and performance. Requires 5+ years experience and strong systems programming skills.
The Senior Cloud Performance Engineer benchmarks and optimizes distributed database and cloud infrastructure performance while developing chaos engineering tools and initiatives. The role requires 6+ years of experience with scalable distributed systems, programming in Go, C/C++, or Java, Kubernetes, and a major public cloud provider.
Leads performance benchmarking, optimization, capacity planning, and chaos engineering for large-scale distributed cloud database systems. Requires 6+ years of software development experience, strong cloud infrastructure expertise, and proficiency in systems programming and production debugging.
Build tools and practices for benchmarking, optimizing, and testing the resilience of large-scale distributed cloud database systems. The role requires 6+ years of software development experience, strong distributed-systems expertise, cloud infrastructure knowledge, and production debugging skills.
Senior cloud performance engineer responsible for benchmarking and optimizing distributed database and cloud infrastructure systems, while building chaos-engineering tools to improve resilience and scalability. Requires 6+ years of software development experience and expertise in distributed systems and public-cloud infrastructure.
Executes physical deployment of data center whitespace infrastructure including racks, power systems, containment, and cabling. Oversees contractors, ensures compliance with designs and safety standards, and manages handoffs to operations. Requires 5+ years in data center construction and engineering degree.
Leads electrical design, optimization, and roadmap for prefabricated modular AI data centers (Crusoe Spark). Requires 5+ years in modular electrical systems, power distribution for AI compute, and cross-functional collaboration. In-office role in Denver with 10-20% travel.
Builds and scales core infrastructure including Kubernetes clusters and cloud resources to power AI data platform. Collaborates with teams on foundational systems, optimizes for cost/performance, and ensures security/compliance. Requires 5+ years infra experience.
Leads infrastructure evolution, CI/CD, observability, and reliability for scaling platform. Requires 10+ years SRE/infra experience, AWS expertise, and systems thinking. Onsite in NYC.
Leads infrastructure reliability for large-scale GPU clusters, architects systems for training/inference scaling, and builds a high-performing engineering team. Requires deep Linux/distributed systems expertise, GPU production experience, and Kubernetes fluency.
Build and scale enterprise GenAI infrastructure across multi-cloud providers (AWS, Azure, GCP), implementing integrations and architecting systems for regulated industries. Requires 4+ years experience, proficiency in Python/JS/SQL, Kubernetes, and AI technologies like LLMs.
Build and maintain infrastructure tooling for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning to support AI workloads at scale. Requires 3+ years managing large server fleets, strong Python and deep Linux expertise.
Seasoned SRE owning reliability of Kubernetes-based production infrastructure at scale for a generative AI platform. Responsibilities include operating clusters, CI/CD, SLOs, monitoring, automation with AI, and driving improvements via chaos engineering. Requires 5+ years production experience with deep Kubernetes and observability expertise.
Designs, deploys, and manages HashiCorp Vault clusters for secure secret management in on-premises and cloud (AWS/GCP) hybrid environments with Kubernetes integration. Requires 3+ years experience, zero trust principles, IaC tools like Terraform, and automation scripting.
Develops and maintains DataOps platform using Kubernetes, cloud services, and data processing tools to support NCBI developers. Requires strong coding skills, cloud experience, and Linux proficiency; bonus for GitOps and observability tools.
Leads design, implementation, and automation of high-performance networks using Arista EOS, CloudVision, F5 BIG-IP, and protocols like BGP, VxLAN. Requires 10+ years experience, Python scripting, and team leadership for NIH scientific mission.
Designs and deploys high-performance network architectures using Spine-and-Leaf, Arista EOS/CloudVision, and F5 BIG-IP. Automates operations with Python, modernizes legacy networks for cloud, and leads technical efforts requiring 10+ years experience and expertise in BGP, OSPF, VxLAN, EVPN.
Senior Linux System Administrator manages Linux infrastructure, automates tasks with scripting (Python, Bash, etc.), uses Puppet for configuration, troubleshoots complex issues, and supports developers at NCBI. Requires 7+ years Linux admin experience and strong automation skills.
Builds and maintains Red Hat OpenShift clusters on GCP for NCBI's DevOps platform. Supports developers with Kubernetes, GitLab CI/CD, ArgoCD, service mesh, and on-call troubleshooting in a hybrid environment.
Senior Platform DevEx Engineer builds and maintains developer tools including GitLab CI pipelines, Python apps, and Kubernetes templates for NCBI's platform services team. Requires 7+ years experience, strong coding skills, Linux knowledge, and CI/CD familiarity.
Designs and optimizes cloud-based database architectures for petabyte-scale biomedical datasets, implementing partitioning, sharding, and query tuning. Requires 5+ years experience with SQL, cloud services, Linux scripting, and production database operations.
Cloud Engineer supports developers and production services in AWS and Google Cloud environments, managing resources like compute instances, VPCs, security groups, and automation scripting while maintaining stable cloud operations onsite in Bethesda, MD.
Builds and maintains DataOps platform using Kubernetes, cloud services, and data processing tools to support NCBI developers. Requires strong coding, cloud experience, and data infrastructure knowledge; B.S. in STEM preferred.
Leads operations support team resolving deployment issues, debugging microservices, and maintaining SLOs in a GitLab/Kubernetes DevOps platform for NCBI developers. Requires BS in STEM, Linux skills, scripting, and strong communication for on-call support and training.
Builds and maintains developer experience tools including GitLab CI pipelines, Python applications, and Kubernetes templates for NCBI's platform services team. Requires coding fluency in Python/C++/JS, Linux knowledge, and CI/CD familiarity; supports deployments across languages and environments.
Builds and maintains developer experience platform tools including GitLab CI pipelines, Kubernetes configurations, and Python applications to enable consistent builds, deployments, and monitoring for NCBI's software teams. Requires strong coding skills in languages like Python/C++, Linux familiarity, and CI/CD experience.
Leads platform operations team supporting developers on GitLab and Kubernetes-based DevOps platform. Resolves deployment issues, manages on-call support, trains team members, and ensures SLOs in microservices environment. Requires BS in STEM, Linux skills, and scripting experience.
The Staff DevOps/SRE Engineer will define infrastructure strategy and SRE practices while building reliable, scalable systems for distributed, multi-cloud AI workloads. The role requires 8+ years of experience, deep Kubernetes and infrastructure-as-code expertise, and strong capabilities in automation, observability, and incident response.
Build and scale infrastructure for high-availability, low-latency tax compliance platform handling millions of global transactions. Manage Kubernetes clusters, Postgres at scale, and cloud-native services while participating in on-call and collaborating with customers.
Software Engineer building infrastructure for data platforms, simulation, or technical services in autonomous driving technology. Requires 2+ years experience, strong Python/C++/Go skills, and expertise in distributed systems or related areas.
Build observability infrastructure and AI-powered tools for OpenAI's large-scale production systems, including logging, metrics, and debugging UIs. Requires experience with distributed systems, Kubernetes, AWS, and observability tools.
Build and operate the cloud and edge infrastructure supporting a large IoT product fleet, including AWS resources, device management, CI/CD, observability, and automation. The role requires at least three years of infrastructure and programming experience, with AWS, Kubernetes, and IoT expertise.
Leads proactive reliability engineering and incident-management improvements for a large-scale, multi-cloud streaming platform. The role requires 10+ years in SRE, incident management, or reliability engineering, plus expertise in distributed systems, observability, Kubernetes, cloud infrastructure, and incident tooling.
Designs, deploys, and manages bare-metal on-prem infrastructure including servers, networking, observability, and backend services for a semiconductor fab. Requires hands-on SRE experience, systems programming in Rust/Go/Python, and BS in CS/CE or equivalent.
Builds and scales multi-cloud infrastructure for enterprise AI Agentic workflows, focusing on security, compliance, observability, and developer tools. Requires 5+ years experience with modern infra practices, cloud providers, and languages like Python.
Performance Engineer optimizes throughput and robustness of large-scale ML distributed systems by solving novel performance issues. Requires significant software engineering experience at supercomputing scale and interest in ML.
Owns end-to-end technical lifecycle of mission-critical federal deployments, ensures compliance with security controls, contributes to core codebase, and shapes product direction from field learnings. Requires 2+ years software engineering experience and federal compliance expertise.
Senior Platform Engineer owns and evolves infrastructure for reliability, performance, and cost optimization at scale. Partners with engineers on debugging, observability (Prometheus, Grafana), deployment pipelines (Kubernetes, Terraform), and on-call incident response. Requires 5+ years experience including DevOps/SRE.
Administers and optimizes Atlassian Cloud tools like Jira and Confluence for enterprise collaboration, ITSM, and reporting. Customizes workflows, drives AI initiatives with Rovo, and ensures security/integrations. Requires 3+ years experience.
Senior Software Engineer building Decagon's internal developer platform, focusing on CI/CD pipelines, developer tooling, observability standards, and workflows that accelerate engineering productivity and reduce toil. Requires 4+ years experience in platform/devtools/infra with strong coding and collaboration skills.