Skip to content
xAIxAI

Software Engineer - Network Software and Services

Build scalable software, automation, and frameworks for managing large AI network fabrics, including metrics, provisioning, monitoring, configuration, and remediation. The role requires deep networking expertise and a track record of designing reliable systems that orchestrate large device fleets.

About the job

Responsibilities

  • Build software and tools with extensive metrics coverage for large GPU supercomputing network fabrics used for AI training and customer inference queries.
  • Implement infrastructure-as-code best practices.
  • Enhance deployment pipelines.
  • Ensure robust and secure service delivery across production environments.

Requirements

  • Deep experience collaborating daily with network engineers.
  • Extensive knowledge of physical and logical network topologies and network protocols.
  • Expert knowledge and a proven history of designing scalable and reliable software from the ground up.
  • Experience building and orchestrating tens of thousands of network devices at high speed.
  • Ability to thrive in ambiguity and create metrics that help prioritize team and individual focus.

Compensation & Benefits

  • €80,000–€150,000 base salary.
  • Equity.
  • Comprehensive medical, vision, and dental coverage.
  • Access to a 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance.
  • Various discounts and perks.

Skills

Network Topologies, Network Protocols, Infrastructure As Code, Deployment Pipelines, Metrics Collection, Network Monitoring, Zero-Touch Provisioning, Auto-Remediation, Gpu Supercomputing, Software Orchestration

Grafana Labs

Grafana Labs

United Kingdom
Software Engineer - Platform Metal
£72k+/yrRemoteDevOps / SRE

Build and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.

PostHog

PostHog

Remote

ClickHouse Operations Engineer
No salary listedRemoteDevOps / SRE

Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.