Senior Linux Infrastructure Engineer
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
About the job
Responsibilities
- Diagnose and tune Linux systems under production load, including CPU scheduling, NUMA, memory, page cache, disk and filesystem I/O, and network performance.
- Lead incident troubleshooting using hypothesis-driven investigation; write postmortems and track preventative follow-up work.
- Build and maintain reviewed, tested, version-controlled configuration management code with Salt, Ansible, Chef, or Puppet.
- Build, deploy, and operate containerized workloads on Kubernetes or Nomad.
- Automate operational work with Bash and Python.
- Operate DNS and DHCP services, including zones, resolvers, TTLs, scopes, reservations, and relays.
- Configure and troubleshoot Nginx and HAProxy, including routing, TLS termination, health checks, and connection handling.
- Support Redis and RabbitMQ in production.
- Provision and maintain virtual machines using VMware, Xen, or KVM, including capacity planning, host maintenance, and live migration.
- Manage Vault secrets, dynamic credentials, policies, and rotation.
- Maintain Elastic Stack log aggregation and Nagios, CheckMK, or Icinga alerting.
- Follow Git-based change management and peer review practices; create runbooks, design proposals, and incident write-ups.
Requirements
- 6+ years of experience in Linux systems, infrastructure, or SRE roles.
- Expert Linux knowledge, including performance analysis with tools such as perf, strace, ss, iostat, or bpftrace.
- Strong understanding of processes and signals, systemd, cgroups, namespaces, filesystems, memory exhaustion, and file descriptors.
- Production-scale declarative configuration management experience with Salt, Ansible, Chef, or Puppet.
- Production containerization and orchestration experience with Kubernetes or Nomad.
- Strong Bash and working Python skills.
- Networking fundamentals across VLANs, routing, DNS, DHCP, TCP, and TLS.
- Proxy and load-balancing experience with Nginx or HAProxy.
- Virtualization experience with VMware, Xen, or KVM.
- Daily experience with Git and peer review.
- Excellent written communication and ability to become productive quickly in unfamiliar areas.
Nice-to-haves
- Primary production ownership of Redis or RabbitMQ.
- Vault administration, including policies, authentication methods, and rotation at scale.
- Elastic Stack operations at volume, including index lifecycle, mappings, and cluster tuning.
- Experience in regulated environments or systems where downtime has direct revenue impact.
- Colocation or bare-metal experience, including hardware lifecycle, remote hands, and finite-footprint capacity planning.
Compensation and Benefits
- Base salary: $140,000–$180,000 annually.
- Discretionary performance bonus: 10–12% of base salary.
- Stock purchase options.
- Medical, vision, and dental benefits.
- 401(k) plan.
- Paid vacation, sick time, gym membership reimbursement, commuter benefits, pet insurance, wellness and mental health programs, charitable donation matching, and paid volunteer days.
- Catered lunches, office kitchen, in-building gym, and Metra shuttle when working from the office.
Skills
Linux, Salt, Ansible, Chef, Puppet, Kubernetes, Nomad, Bash, Python, DNS, Dhcp, Nginx, Haproxy, VMware, Vault
Similar jobs
DevOps / SRE jobsDesigns, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.