Operations Support Engineer - Lead
Leads operations support team resolving deployment issues, debugging microservices, and maintaining SLOs in a GitLab/Kubernetes DevOps platform for NCBI developers. Requires BS in STEM, Linux skills, scripting, and strong communication for on-call support and training.
About the job
Duties & Responsibilities
- Identify and resolve operational problems in a micro-service environment
- Work with developers to resolve deployment and runtime problems
- Perform analysis and debugging work across multiple technologies
- Prioritize issues to keep applications within error budgets and meeting their SLOs
- Provide technical solutions to a wide range of problems and user requests
- Document process, procedures and SOPs by soliciting feedback and suggestions from team members
- Compile postmortems and action items to minimize future outages
- Interview other people for team member roles, and decide which ones to recommend for hire
- Train new team members, and assist them with issues
- Provide on-call support to NCBI's internal developers and other staff
Requirements
- BS degree in STEM or equivalent experience
- Customer-focused, team-oriented disposition
- Good systems debugging skills
- Comfortable with the Linux environment or UNIX CLI
- Experience with some programming or scripting language
- Have experience creating processes, procedures and SOP documentation
- General understanding of TCP/IP, HTTP, and related protocols
- Initiative to take ownership of tasks and drive them to completion
- Comfortable dealing with users with varying levels of IT knowledge
- Eager to learn new technologies
- Strong communication and soft skills to interface with customers, peers and management
- Good judgement, sense of integrity, and responsibility
Preferred Experience / Skillsets
- Kubernetes, OpenShift, Cloud or Linux experience
- Experience with:
- Service Reliability Engineering in any capacity
- Linux systems administration
- Automated CI servers, especially TeamCity and/or GitLab
- Automation programming/scripting in any of: bash, Ruby, Python, Go, Java, Scala, Rust, C++, Perl
- Automated configuration management, such as Puppet, Ansible, Chef, bcfg2, cfengine, etc. (Puppet is preferred)
- Version control systems, especially git
- Service Mesh technologies (e.g., linkerd, Istio)
- Configuring or using monitoring and alerting technologies (TIGK stack, Grafana, Prometheus, OpsGenie)
- Confluence, Jira, and Microsoft Office suite
- GitOps tools, especially ArgoCD
- Google Anthos
- Understanding of:
- Linux internals (system calls, file systems, processes, etc.)
- Linux network configuration
- Linux application containerization, especially Docker
- Attached network storage technologies
- Cloud computing environment such as AWS, GCP or Azure
- Automated CI/CD pipelines
- Distributed systems design principles
Skills
Kubernetes, Linux, GitLab, Docker, Python, Bash, Puppet, Ansible, Prometheus, Grafana, AWS, GCP, Istio, Argo CD, Jira
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.