Senior Software Engineer
Senior software engineer owning a multi-environment compute platform for simulation orchestration and workload scheduling at scale. Requires 5+ years backend/distributed systems experience and deep Kubernetes expertise.
About the job
Responsibilities
- Develop a multi-environment compute platform for workload scheduling, resource allocation, and node lifecycle.
- Collaborate with customers and internal teams to translate compute needs into platform features.
- Solve distributed systems challenges including GPU scheduling, autoscaling, and resource efficiency.
- Ensure high reliability and scalability for production simulation systems.
- Drive architecture decisions to expand simulation orchestration into a company-wide compute platform.
Requirements
- 5+ years of backend engineering experience, with a track record of owning production distributed systems.
- 4+ years of coding in Golang, Python, Java, or C++.
- 4+ years building complex backend systems on major cloud providers.
- Deep Kubernetes expertise, including pod scheduling, node lifecycle, and autoscaling under load.
- Experience building or operating workload scheduling systems (e.g., dispatch, bin-packing, quota enforcement).
- Self-starter with experience driving production system development from 0 to 1.
Nice-to-Haves
- Experience scheduling GPU workloads or managing GPU capacity at scale.
- Multi-cloud (AWS, OCI, GCP, Azure) or on-prem Kubernetes deployments.
- Experience managing multi-cluster environments using Infrastructure as Code.
- Experience building custom Kubernetes controllers or operators.
- Experience with cloud cost optimization strategies for high-scale compute workloads.
Compensation
- Base salary range: $153000 - $220000 USD annually.
- Total compensation package may also include equity, comprehensive health, dental, vision, life and disability insurance coverage, 401k retirement benefits with employer match, learning and wellness stipends, and paid time off.
Skills
Go, Python, Java, C++, Kubernetes, AWS, GCP, Azure, Oci, Infrastructure As Code
Similar jobs
DevOps / SRE jobsBuild and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.