Latest DevOps / SRE jobs
Job results
Maintains and expands data centers supporting AI/ML infrastructure by installing, troubleshooting, and repairing servers, networks, and hardware. Requires 3-5 years data center experience, Linux knowledge, physical lifting ability, and on-call availability.
Builds tools and systems to enhance developer productivity, including build infrastructure, CI pipelines, service APIs, and supply chain security. Requires 2+ years experience in Java, Go, TypeScript or similar, with strong problem-solving skills.
Builds and leads development of Palantir's high-scale observability platform, handling log/metric/trace ingestion, processing, monitoring, and alerting. Requires 5+ years software experience, strong Go/Java skills, system design expertise, and mentorship abilities.
Network Engineer develops and operates a high-performance blockchain network of 1000 nodes, optimizing for security, scalability, and 100M requests/second. Requires 3-5 years experience specializing in cloud ops, optimizations, kernel, and strong communication.
Develops and operates scalable, secure blockchain infrastructure, managing 1000-node networks and optimizing high-throughput systems. Requires 5-10 years experience in low-level systems programming with kernel/data structures expertise and startup background.
Leads response to critical product outages by triaging, troubleshooting, and coordinating resolutions across teams. Requires technical background in CS/Engineering, strong problem-solving, and comfort with 24/7 on-call in fast-paced environments.
Builds and maintains internal platform infrastructure including Kubernetes clusters, stateful services like databases and Bitcoin/LN nodes, and observability tools. Advises dev teams on integrations; requires strong Linux, networking, cloud, and systems programming expertise.
Senior Site Reliability Engineer automates operational processes, manages secure infrastructure across hybrid data centers and cloud, and improves workflows for ML/data teams. Requires 3-5 years experience with Linux, containers, and automation tools.
Designs and builds infrastructure tools for Lightning Network including automated channel management, liquidity optimization algorithms, monitoring systems, and network health metrics. Requires systems programming expertise in Go/C/C++, Bitcoin knowledge, and secure scalable systems experience.
Systems Engineer/DevOps role focused on automating operations, managing hybrid on-prem data centers and AWS infrastructure, and ensuring reliability for ML/SaaS services. Requires 1-2 years experience with Linux, containers, and automation tools.
Develops and maintains a Linux-based platform for security instrumentation, processing data into databases and publishing for distributed systems. Requires expertise in C/Go/Bash, Linux, security practices, and CI/CD.
The Senior SRE will design, automate, and operate highly available infrastructure for a real-time SaaS platform handling massive datasets. The role combines reliability engineering, observability, incident response, security collaboration, and software development.
Senior SRE/DevOps Engineer owns and operates AWS infrastructure and Kubernetes-based application stacks for Metabase Cloud, debugs issues, builds automation tooling, and improves deployments. Requires 5+ years experience with strong Kubernetes, AWS, Terraform, and modern languages like Python/Go.
Site Reliability Engineer builds, operates, and maintains scalable infrastructure for air-gapped production environments, focusing on Linux servers, cloud/on-prem systems, automation, and troubleshooting. Requires 4+ years Linux admin experience, active security clearance, and proficiency in programming/scripting.
Site Reliability Operations Analyst streamlines workflows, stabilizes projects, and resolves issues for Palantir deployments, primarily onsite with US Government customers. Requires 3+ years project management experience, US security clearance eligibility, and 25-75% travel.
Designs and builds internal compute infrastructure platforms using Palantir products and open-source tools. Requires 3+ years in software development on core infrastructure, expertise in Go/Python/Rust, containers, Kubernetes, and cloud providers.