Skip to content

Platform Ops Lead

Leads platform operations team supporting developers on GitLab and Kubernetes-based DevOps platform. Resolves deployment issues, manages on-call support, trains team members, and ensures SLOs in microservices environment. Requires BS in STEM, Linux skills, and scripting experience.

About the job

Duties and Responsibilities

  • Identify and resolve operational problems in a micro-service environment
  • Work with developers to resolve deployment and runtime problems
  • Perform analysis and debugging work across multiple technologies
  • Prioritize issues to keep applications within error budgets and meeting their SLOs
  • Provide technical solutions to a wide range of problems and user requests
  • Document processes, procedures and SOPs by soliciting feedback and suggestions from team members
  • Compile postmortems and action items to minimize future outages
  • Interview other people for team member roles, and decide which ones to recommend for hire
  • Train new team members, and assist them with issues
  • Provide on-call support to NCBI's internal developers and other staff

Requirements

  • BS degree in STEM or equivalent experience
  • Customer-focused, team-oriented disposition
  • Good systems debugging skills
  • Comfortable with the Linux environment or UNIX CLI
  • Experience with some programming or scripting language
  • Have experience creating processes, procedures and SOP documentation
  • General understanding of TCP/IP, HTTP, and related protocols
  • Initiative to take ownership of tasks and drive them to completion
  • Comfortable dealing with users with varying levels of IT knowledge
  • Eager to learn new technologies
  • Strong communication and soft skills to interface with customers, peers and management
  • Good judgement, sense of integrity, and responsibility

Preferred Experience and Skillsets

  • Kubernetes, OpenShift, Cloud or Linux experience

  • Experience with:

    • Service Reliability Engineering in any capacity
    • Linux systems administration
    • Automated CI servers, especially TeamCity and/or GitLab
    • Automation programming/scripting in any of: bash, Ruby, Python, Go, Java, Scala, Rust, C++, Perl
    • Automated configuration management, such as Puppet, Ansible, Chef, bcfg2, cfengine, etc. (Puppet is preferred)
    • Version control systems, especially git
    • Service Mesh technologies (e.g., linkerd, Istio)
    • Configuring or using monitoring and alerting technologies (TIGK stack, Grafana, Prometheus, OpsGenie)
    • Confluence, Jira, and Microsoft Office suite
    • GitOps tools, especially ArgoCD
    • Google Anthos
  • Understanding of:

    • Linux internals (system calls, file systems, processes, etc.)
    • Linux network configuration
    • Linux application containerization, especially Docker
    • Attached network storage technologies
    • Cloud computing environment such as AWS, GCP or Azure
    • Automated CI/CD pipelines
    • Distributed systems design principles

Benefits and Salary

  • Competitive benefits package that includes medical, dental and vision coverage, 401k plan with employer contribution, paid holidays, vacation, and tuition reimbursement
  • Competitive salary commensurate with experience and location. The targeted range for this position is $135,000 - $165,000

Skills

Kubernetes, GitLab, Linux, Docker, Ansible, Puppet, Prometheus, Grafana, Python, Bash, CI/CD, GitOps, Argo CD, Istio, AWS

Upstart

Upstart

United States

Senior DevOps Engineer
$136k+/yrRemote3+ YOEDevOps / SRE

Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

tastytrade

tastytrade

Chicago, IL

Senior Linux Infrastructure Engineer
$140k+/yrHybrid6+ YOEDevOps / SRE

Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Okta

Okta

San Francisco, CA

Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.