Linux Systems Engineer (USA)
Hands-on Linux Systems Engineer builds and maintains bare-metal servers, manages storage like ZFS, automates with Ansible and Bash, and ensures production reliability. Requires 3+ years Linux experience, physical server management, and on-call rotation with data center travel.
About the job
Responsibilities
- Build, provision, and maintain bare-metal Linux servers, including OS installation, configuration, and lifecycle management
- Own server infrastructure from hardware through operating system and core services, ensuring stability and performance
- Configure and manage storage systems, including ZFS and enterprise storage platforms (e.g., NetApp, Dell, or similar)
- Monitor system health and performance; troubleshoot issues and implement durable, long-term fixes
- Develop and maintain automation using Ansible and Bash to standardize provisioning, configuration, and operations
- Perform system patching, upgrades, and capacity planning across a growing server fleet
- Participate in incident response, root cause analysis, and continuous improvement of system reliability
- Collaborate with engineering and infrastructure teams to support application performance on Linux systems
- Contribute to documentation, runbooks, and operational best practices
- Support data center operations as needed, including hardware troubleshooting, racking, cabling, and server replacements
- Travel to data centers periodically for maintenance, expansions, and issue resolution
Requirements
- Bachelor’s degree in a technical field or equivalent practical experience
- 3+ years of experience managing Linux systems in a production environment
- Strong expertise in Linux at the system level, including OS installation, configuration, and troubleshooting on bare metal
- Proven experience provisioning and managing physical servers at scale (100+ servers preferred)
- Hands-on experience with storage systems, including ZFS; familiarity with enterprise storage vendors (NetApp, Dell, or similar) is strongly preferred
- Proficiency with Ansible for configuration management and automation (required)
- Strong Bash scripting skills; Python is a plus but not required
- Solid understanding of system performance, resource management, and reliability in production environments
- Experience working in environments where uptime, precision, and operational discipline are critical
- Willingness to perform occasional hands-on hardware work and travel to data centers as needed
- Ability to participate in an on-call rotation
Compensation and Benefits
- Base salary: $130,000 - $150,000, determined by education and experience
- Performance-based bonus
- PPO health, dental, and vision insurance fully covered for employees and dependents
- Pre-tax commuter benefits
- Weekly company-sponsored meals
Skills
Linux, Ansible, Bash, Zfs, Netapp, Dell, Python
Similar jobs
DevOps / SRE jobsBuilds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.
Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.