Skip to content
AxleAxle

IT Operations Technical Lead

Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.

About the job

Responsibilities

  • Lead IT operations aligned with ITIL processes, including incident, problem, change, and release management.
  • Manage Linux and Windows environments across hybrid cloud and on-premises infrastructure.
  • Drive incident response, root-cause analysis, and service restoration for mission-critical systems.
  • Design and maintain golden images, patching strategies, and system-hardening standards.
  • Lead patch management and vulnerability remediation programs.
  • Develop automation solutions, including AI-assisted development approaches, to improve operational efficiency.
  • Support and optimize infrastructure for AI/ML workloads, including provisioning, scaling, and performance tuning.
  • Manage GPU-enabled environments for high-performance computing and machine-learning use cases.
  • Optimize monitoring, logging, alerting, and observability frameworks.
  • Manage and mentor systems engineers and provide technical guidance and performance oversight.
  • Collaborate with architecture, security, and development teams to improve reliability and scalability.
  • Support hybrid cloud and on-premises data-center environments.
  • Maintain documentation, runbooks, standard operating procedures, and operational-readiness materials.

Requirements

  • 5+ years leading operations teams and driving operational process and technology improvements.
  • Experience implementing and operating within ITIL frameworks.
  • 10+ years of hands-on Unix/Linux experience, including CentOS and Red Hat systems administration in large-scale distributed environments.
  • Experience with incident management, patching, system hardening, and production support.
  • Experience building and maintaining golden images and standardized environments.
  • Strong scripting and automation skills.
  • Experience with configuration management and automation tools.
  • Strong understanding of networking fundamentals, including DNS, TCP/IP, firewalls, and load balancing.
  • Experience with monitoring and logging tools.
  • Cloud build-out or migration experience with AWS, Google Cloud, or Microsoft Azure.
  • 2+ years of experience with CI/CD and automation tools.
  • Experience supporting AI/ML workloads or data-intensive platforms.
  • Familiarity with GPU-based compute environments.
  • Knowledge of security best practices and compliance frameworks.
  • Willingness to learn and adopt emerging technologies.

Nice-to-Haves

  • Knowledge of NIST 800-53, FedRAMP, or FISMA.
  • ITIL, Linux, AWS, Azure, or Kubernetes certifications, including CKA/CKAD.
  • CCNA or CCNP certification.

Compensation and Benefits

  • Base salary: $150,000–$170,000 USD.
  • Medical, dental, and vision coverage for employees.
  • Paid time off and holidays.
  • 401(k) match up to 5%.
  • Educational benefits.
  • Employee referral bonus.
  • Flexible spending and transportation reimbursement accounts.

Skills

Linux, Windows Server, Itil, Python, Bash, PowerShell, Ansible, Terraform, Puppet, Chef, AWS, GCP, Microsoft Azure, Kubernetes, Prometheus

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

PointClickCare

PointClickCare

United States

Senior Site Reliability Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.

Clear Street

Clear Street

New York, NY

Senior Software Engineer - Platform Engineer
$150k+/yrHybrid5+ YOEDevOps / SRE

Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.

Okta

Okta

San Francisco, CA

Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.