What You’ll Be Working On
- Tuesday - Saturday Shift Pattern. Tuesday - Friday on site - Saturday Remote.
- Customer Support: Provide exceptional technical support to customers via Zendesk, meeting SLAs and maintaining high CSAT (95%+).
- On-Call Rotation: Participate in a 24/7 on-call rotation to ensure timely resolution of critical issues.
- Incident Management: Primary point of contact for incident management, focusing on initial triage, communication, and procedural rigor throughout the incident lifecycle. Lead response efforts, ensuring clear communication with both technical and non-technical stakeholders and acting as a customer advocate to minimize disruption.
- Troubleshooting: Diagnose and resolve issues related to VMs, hardware failures, and scaling tests using CLI and internal tools.
- Alert Triage and Maintenance: Manage alert triage, prepare for maintenance windows, and conduct node delivery testing.
- Collaboration: Work closely with SRE, Networking, and Storage teams from initial triage to root cause analysis (RCA) delivery.
- Global Teamwork: Adhere to global team collaboration and handoff processes for ticketing and on-call procedures.
- Knowledge Sharing: Develop onboarding/training materials, knowledge base documentation, and standard operating procedures (SOPs).
What You’ll Bring to the Team
- Education/Experience: Bachelor's degree in IT, Computer Science, Engineering, or a related field, or 4+ years of equivalent technical experience.
- Linux Proficiency: Strong command-line interface (CLI) skills in Linux environments.
- Version Control: Proficiency with Git for code management and collaboration.
- Customer Support Experience: 5+ years of experience in a customer support role, ideally within cloud, storage, or networking environments.
- Cloud Technologies: Experience with container orchestration (e.g., Kubernetes), workload management (e.g., Slurm, Terraform), and monitoring tools (e.g., Grafana).
- Public Cloud Knowledge: Familiarity with other public cloud platforms (e.g., AWS, Azure, GCP).
- Communication Skills: Excellent communication and customer service skills, including the ability to prioritize competing escalations.
- HPC Knowledge: Understanding of HPC technologies such as Infiniband, RDMA, RoCE, and Software Defined Networking (SDN).
Bonus Points
- Certifications: CKA, CKAD, CKS, KCNA, AWS Machine Learning - Specialty, Data Analytics - Specialty, Solutions Architect - Professional, Developer - Associate, NVIDIA AI Infrastructure and Operations, Generative AI and LLMs, Generative AI Multi-modal, Infiniband, Linux Foundation IT Associate, System Administrator.
- Cloud Expertise: Deep understanding of specific cloud platforms and services.
- Automation Skills: Experience with automation tools and scripting languages.
- Problem-Solving Abilities: Demonstrated ability to analyze complex technical issues and develop effective solutions.
- Collaboration and Mentorship: Proven ability to mentor, train, and onboard colleagues.
- Passion for Sustainability: A strong interest in contributing to a more sustainable future through technology.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
Compensation Range
Compensation will be paid in the range of up to $145,000 - $175,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.