Software Engineer - Network Software and Services
Build scalable software, automation, and frameworks for managing large AI network fabrics, including metrics, provisioning, monitoring, configuration, and remediation. The role requires deep networking expertise and a track record of designing reliable systems that orchestrate large device fleets.
About the job
Responsibilities
- Build software and tools with extensive metrics coverage for large GPU supercomputing network fabrics used for AI training and customer inference queries.
- Implement infrastructure-as-code best practices.
- Enhance deployment pipelines.
- Ensure robust and secure service delivery across production environments.
Requirements
- Deep experience collaborating daily with network engineers.
- Extensive knowledge of physical and logical network topologies and network protocols.
- Expert knowledge and a proven history of designing scalable and reliable software from the ground up.
- Experience building and orchestrating tens of thousands of network devices at high speed.
- Ability to thrive in ambiguity and create metrics that help prioritize team and individual focus.
Compensation & Benefits
- €80,000–€150,000 base salary.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance.
- Various discounts and perks.
Skills
Network Topologies, Network Protocols, Infrastructure As Code, Deployment Pipelines, Metrics Collection, Network Monitoring, Zero-Touch Provisioning, Auto-Remediation, Gpu Supercomputing, Software Orchestration
Similar jobs
DevOps / SRE jobsBuild and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.