IT Operations Technical Lead
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
About the job
Responsibilities
- Lead IT operations aligned with ITIL processes, including incident, problem, change, and release management.
- Manage Linux and Windows environments across hybrid cloud and on-premises infrastructure.
- Drive incident response, root-cause analysis, and service restoration for mission-critical systems.
- Design and maintain golden images, patching strategies, and system-hardening standards.
- Lead patch management and vulnerability remediation programs.
- Develop automation solutions, including AI-assisted development approaches, to improve operational efficiency.
- Support and optimize infrastructure for AI/ML workloads, including provisioning, scaling, and performance tuning.
- Manage GPU-enabled environments for high-performance computing and machine-learning use cases.
- Optimize monitoring, logging, alerting, and observability frameworks.
- Manage and mentor systems engineers and provide technical guidance and performance oversight.
- Collaborate with architecture, security, and development teams to improve reliability and scalability.
- Support hybrid cloud and on-premises data-center environments.
- Maintain documentation, runbooks, standard operating procedures, and operational-readiness materials.
Requirements
- 5+ years leading operations teams and driving operational process and technology improvements.
- Experience implementing and operating within ITIL frameworks.
- 10+ years of hands-on Unix/Linux experience, including CentOS and Red Hat systems administration in large-scale distributed environments.
- Experience with incident management, patching, system hardening, and production support.
- Experience building and maintaining golden images and standardized environments.
- Strong scripting and automation skills.
- Experience with configuration management and automation tools.
- Strong understanding of networking fundamentals, including DNS, TCP/IP, firewalls, and load balancing.
- Experience with monitoring and logging tools.
- Cloud build-out or migration experience with AWS, Google Cloud, or Microsoft Azure.
- 2+ years of experience with CI/CD and automation tools.
- Experience supporting AI/ML workloads or data-intensive platforms.
- Familiarity with GPU-based compute environments.
- Knowledge of security best practices and compliance frameworks.
- Willingness to learn and adopt emerging technologies.
Nice-to-Haves
- Knowledge of NIST 800-53, FedRAMP, or FISMA.
- ITIL, Linux, AWS, Azure, or Kubernetes certifications, including CKA/CKAD.
- CCNA or CCNP certification.
Compensation and Benefits
- Base salary: $150,000–$170,000 USD.
- Medical, dental, and vision coverage for employees.
- Paid time off and holidays.
- 401(k) match up to 5%.
- Educational benefits.
- Employee referral bonus.
- Flexible spending and transportation reimbursement accounts.
Skills
Linux, Windows Server, Itil, Python, Bash, PowerShell, Ansible, Terraform, Puppet, Chef, AWS, GCP, Microsoft Azure, Kubernetes, Prometheus
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.
Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.