Senior Linux Infrastructure Engineer
Builds, automates, and maintains scalable, secure Linux infrastructure using modern tooling. Requires 7+ years Linux sysadmin experience, deep OS/network knowledge, Python/Ruby scripting, config management (Ansible/Puppet/Chef), and monitoring (Prometheus/Grafana).
About the job
Responsibilities
- Developing operational tooling in Python, Ruby, or similar language (beyond shell scripts for API integrations, data processing, and system orchestration tasks)
- Design and implement systems that are highly available, self healing, scalable, and performant
- Create comprehensive metrics, monitoring, visualizations, and alerting systems applied to all parts of the infrastructure
- Design and implement all systems components with “Security First” principles
- Heavily engage with AI tools for all aspects of development and support
- Participate in an on-call rotation
Requirements
- At least 7 years of working as a Linux Systems Administrator or Engineer
- In-Depth understanding of Linux fundamentals regarding filesystems, process scheduling, virtual memory management, and other base primitives
- In-Depth understanding of the TCP/IP network protocol and addressing
- Significant experience developing operational tooling in Python, Ruby, or similar language
- Significant experience with Configuration Management platform (Ansible, Puppet, Chef)
- Experience with metrics, monitoring and visualization tools (Prometheus, Influx, Grafana)
Preferred Qualifications
- Experience with CI/CD tools (Jenkins, Github Actions)
- Experience with containers and container orchestration
- Experience with distributed storage (Ceph, Lustre, GPFS)
- Experience with RDBMS (Postgres, MySQL)
Skills
Linux, Python, Ruby, Ansible, Puppet, Chef, Prometheus, Influx, Grafana, TCP/IP, CI/CD, Jenkins, GitHub Actions, Containers, Kubernetes Or Container Orchestration
Similar jobs
DevOps / SRE jobsSenior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.