Staff Production Engineer
The Staff Production Engineer will operate and improve reliable production infrastructure, with a strong focus on data center networking, capacity planning, troubleshooting, automation, and hardware operations. The role requires 5+ years of production engineering experience, networking expertise, and proficiency with infrastructure and CI/CD tools.
About the job
Responsibilities
- Analyze network capacity, collaborate on growth plans, and contribute to network builds in data centers and Points of Presence (PoPs).
- Provide Tier 1 troubleshooting for switch provisioning, server build failures caused by network issues, and complex physical-layer network defects.
- Manage console DNS mapping, troubleshoot connectivity issues, verify automation scripts, conduct network audits, and ensure seamless handoffs to internal customers.
- Oversee network hardware failures, switch swaps, appliance replacements, and RMA processes.
- Troubleshoot Layer 1 optical-network issues using OTDR tests.
- Partner with data center and network engineering teams on structured cabling for data center expansions and fabric switch uplifts.
- Identify recurring network issues and lead operational excellence projects.
- Conduct network studies, analyze link and hardware stability, and recommend improvements.
- Evaluate and adopt new technologies and methodologies through internal working groups.
Requirements
- 5+ years of professional production engineering experience.
- 5+ years contributing to architecture and design of new and existing systems, including design patterns, reliability, and scaling.
- Bachelor's degree in Computer Science or a related field, or 8+ years of relevant work experience.
- Solid understanding of infrastructure design and operational trade-offs.
- Experience writing high-quality code in Python, Go, or a similar language.
- Experience with Docker, Kubernetes, Ansible, CloudFormation, or Terraform.
- Experience with modern CI/CD practices and build systems such as GitLab CI/CD, CircleCI, or GitHub Actions.
- Experience with logging, monitoring, and alerting systems.
- Experience with Unix/Linux environments, TCP/IP, network programming, and information security best practices.
- Solid understanding of IP subnetting, Layer 2 and Layer 3 networking, VLANs, MAC addresses, port speeds, optics, and routing.
- 2+ years of experience with BGP, MPLS, L3VPN, VPLS, Multicast, CoS, TCP, and IPv4/IPv6.
- 2+ years of experience as a Network Engineer for a content or network provider.
- Experience with data center network architectures, including CLOS.
- Experience with structured cabling in data center environments, including SMF and MMF.
- Experience managing vendors for logistics, infrastructure racking and cabling, and ordering.
- Excellent communication skills and alignment with company values.
Nice-to-Haves
- Bachelor's degree in Computer Science, Engineering, or a related technical discipline.
- 2+ years of programming experience in Python, Perl, or another scripting language.
- Experience with Ansible, REST/gRPC, and NETCONF.
- Network operations experience with a systematic troubleshooting approach.
- CCNP or JNCIS-equivalent certification.
- Experience with open-source Network Operating Systems.
- Interest in engineering and deploying in-house network hardware and software.
- Experience testing and deploying new hardware platforms.
Benefits and Compensation
- Full social security coverage and contributions to provident, trade union, and pension funds, with additional pension options.
- Optional access to Global Life Insurance and private health insurance.
- Generous maternity, paternity, parental, and sick leave policies.
- Compensation is determined by education, experience, knowledge, skills, abilities, internal equity, and market data; no specific salary range is provided.
Skills
Python, Go, Docker, Kubernetes, Ansible, Terraform, GitHub Actions, Linux, TCP/IP, BGP, Mpls, Ipv4/Ipv6, Clos Networking, Netconf, Otdr
Similar jobs
DevOps / SRE jobsLeads the design, development, and operation of Stripe’s large-scale CI and developer productivity systems. Requires 10+ years of hands-on software development, distributed-systems expertise, technical leadership, and mentoring experience.
Leads the technical direction of multi-cloud Kubernetes capacity management and workload placement across Datadog’s large-scale infrastructure. The role requires strong systems programming experience, ideally in Go, cloud infrastructure expertise, and the ability to influence architecture across teams.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.
Staff Site Reliability Engineer responsible for the performance, scalability, deployment robustness, vulnerability management, and incident response of a large-scale cloud data platform. The role requires deep Kubernetes and multi-cloud expertise, infrastructure automation, scripting, Linux administration, and cloud networking experience.