Director, IT Operations
Leads global IT operations and infrastructure for a healthcare SaaS platform, overseeing reliability, help desk services, compliance, cloud optimization, and AI-enhanced operations. Requires 8+ years of relevant experience and 5+ years managing technical teams.
About the job
Responsibilities
- Lead a global team of 10+ resources covering help desk services, facility management, SOC policy enforcement, and asset management.
- Design and oversee infrastructure architecture supporting 99.5%+ uptime in a healthcare SaaS environment.
- Deliver help desk services meeting a 95% response and resolution SLA.
- Co-own incident management, escalation protocols, and post-incident reviews.
- Manage infrastructure spending and cloud resource optimization.
- Lead technical standardization across development, security, and product teams.
- Identify and implement AI/ML tools for predictive alerting, anomaly detection, root-cause analysis, and operational automation.
- Define operational requirements for AI model deployment, monitoring, and governance.
- Establish observability and alerting standards for AI systems at scale.
- Build operational runbooks and playbooks for safe, rapid experimentation.
- Partner with Security on SOC 2, HIPAA, and state healthcare compliance.
- Manage access control, secrets management, and least-privilege architectures.
- Own disaster recovery and business continuity planning.
- Lead quarterly compliance audits and remediation workflows.
- Establish SLAs/SLOs and own service performance.
- Design monitoring, alerting, and observability infrastructure.
- Lead capacity planning and infrastructure roadmaps.
- Communicate infrastructure health, risks, and roadmaps to leadership and cross-functional teams.
- Support product and engineering teams with infrastructure provisioning and troubleshooting.
- Manage vendor and cloud-provider relationships.
- Mentor and develop operational staff.
Required Qualifications
- 8+ years of experience in IT operations, infrastructure engineering, or site reliability.
- 5+ years in a leadership role managing technical teams, including 5+ direct reports.
- Experience operating mission-critical SaaS or healthcare platforms.
- Deep expertise with AWS, Google Cloud, or Azure and infrastructure as code.
- Strong understanding of HIPAA compliance, security baselines, and healthcare regulations.
- Experience designing systems for 99.5%+ uptime and handling production incidents.
- Track record of building automation-first operational cultures, blameless postmortems, and effective on-call practices.
- Ability to build processes from scratch in a fast-growing organization.
Preferred Qualifications
- Experience building IT organizations in early-stage or hypergrowth companies.
- Hands-on experience with Docker, Kubernetes, and modern DevOps stacks.
- Healthcare experience, including EMR/EHR systems or telehealth platforms.
- Experience supporting machine learning or AI infrastructure at scale.
- Familiarity with database operations, replication, and backup/recovery strategies.
- Knowledge of healthcare data architecture and privacy requirements.
- Startup or consulting experience.
- Experience evaluating or implementing AIOps and other AI/ML operations tooling.
- Track record of using AI to improve operational efficiency or decision-making.
Compensation and Benefits
- Medical, dental, and vision plans.
- Flexible spending and health savings accounts.
- Flexible paid time off.
- 401(k) with company match.
- Life insurance, pet insurance, and other benefits.
Skills
AWS, GCP, Azure, Infrastructure As Code, Docker, Kubernetes, AI/ML, Observability, Incident Management, HIPAA, SOC 2, Disaster Recovery, Database Replication, Backup And Recovery, Site Reliability
Similar jobs
IT Support jobsLeads a multidisciplinary IT Operations organization spanning support, logistics, and AV/executive services for a large technology company. The role requires management experience across 20+ employees, substantial IT budget ownership, organizational-change leadership, and familiarity with enterprise tooling and AI-enabled support.
Leads the strategy, delivery, integration, security, and optimization of the enterprise applications portfolio while managing a technical team, vendors, budget, and compliance programs. Requires 8+ years of enterprise applications leadership and 3+ years of people management.
Leads the strategy, architecture, governance, and hands-on evolution of a global People Systems ecosystem centered on Workday. Requires 10+ years of HRIS or People Systems experience, deep Workday expertise, and the ability to lead cross-functional initiatives and vendors.
Leads and scales a hands-on corporate IT function covering support, endpoints, identity, enterprise applications, office technology, and compliance controls. Requires corporate IT leadership experience, strong customer-service execution, and the ability to manage people and projects while remaining technically engaged.
Leads Docker’s global IT organization, setting strategy while overseeing help desk services, enterprise applications, devices, and employee productivity tooling. Requires 12+ years of IT leadership experience, SaaS/cloud expertise, and a bachelor’s degree.