Senior Backend Software Engineer, Cloud Management
Senior backend engineer building foundational Crusoe Cloud services, including the API gateway, resource management, quotas, notifications, and onboarding workflows. Requires distributed-systems experience, strong PostgreSQL skills, modern compiled-language expertise, and familiarity with reliable event-driven and asynchronous systems.
About the job
Responsibilities
- Design, develop, and maintain scalable, reliable services powering Crusoe Cloud’s user-facing experiences.
- Own and evolve the API gateway, including routing, authentication and authorization enforcement, rate limiting, API versioning, tracing, and tenant isolation.
- Build the notification and eventing backbone with reliable delivery and user preference controls.
- Evolve quota and entitlement systems, including self-service quota requests and fast, correct enforcement under contention.
- Improve customer time-to-first-GPU by hardening and automating onboarding, provisioning, and identity flows.
- Extend resource management models, lifecycle state machines, reconciliation loops, and infrastructure offerings.
- Contribute to architectural decisions supporting reliability and maintainability.
- Collaborate with product, design, and cross-functional teams to evaluate tools, frameworks, and customer needs.
- Mentor engineers, improve hiring practices, and contribute to an inclusive engineering culture.
Requirements
- 3+ years of software development experience.
- Experience programming with modern compiled languages such as Go, Rust, Java, or C++; Go is used primarily by the team.
- Experience designing versioned public APIs using gRPC/protobuf or REST, plus related SDK, CLI, or Terraform provider surfaces.
- Knowledge of API gateway and edge concerns, including routing, authentication, authorization, rate limiting, throttling, quota checks, validation, and graceful degradation.
- Experience designing and scaling fault-tolerant distributed systems and managed cloud services.
- Strong PostgreSQL skills, including schema design, transactions, isolation levels, locking, and safe online migrations.
- Experience with message queues or streaming systems such as Kafka, NATS, or Pub/Sub.
- Knowledge of reliable event-driven patterns, including at-least-once delivery, idempotent consumers, retries, backoff, and dead-letter handling.
- Experience building long-running, resumable workflows and reconciliation loops, using state machines, control loops, or workflow engines such as Temporal.
- Strong fundamentals in data structures, algorithms, microservices, and infrastructure tooling.
- Experience with Docker, Kubernetes, Terraform, and CI/CD systems.
- Experience defining SLOs, instrumenting services with metrics, logs, and traces, and debugging production systems while on call.
- Strong cross-functional collaboration, communication, and mentorship skills.
Nice-to-haves
- Experience building infrastructure tooling.
- Experience with cloud customer experience platforms, resource management, quotas, entitlements, notifications, or onboarding systems.
- Experience improving engineering hiring and onboarding processes.
Skills
Go, Rust, Java, C++, gRPC, Protocol Buffers, Rest, Postgres, Kafka, Nats, Pub/Sub, Temporal, Docker, Kubernetes, Terraform
Similar jobs
Backend Engineering jobsBuild and operate high-throughput blockchain infrastructure, APIs, and platform primitives integrating protocols such as Ethereum and Bitcoin with internal services. Requires 5+ years of software engineering experience, distributed-systems expertise, and hands-on crypto infrastructure experience.
Design, build, and operate Cloudflare’s globally distributed cache and reverse-proxy data plane, improving performance, correctness, and resilience across the edge. Requires at least 4 years of production systems experience and proficiency in a systems or backend language.
Senior backend engineer designing and operating reliable billing and financial systems, APIs, data models, and distributed workflows. The role requires 5+ years of professional software development experience, strong backend expertise, and collaboration across Product, Finance, Operations, and Data.
Senior individual contributor responsible for designing, building, operating, and improving large-scale backend services, APIs, and telemetry pipelines in Go and Python. The role requires production systems ownership, distributed-systems expertise, incident leadership, mentoring, and technical design leadership.
Build and operate backend services, data pipelines, storage, and retrieval systems that provide trusted context to agentic platforms and product applications. The role requires 8+ years of software engineering experience, distributed-systems expertise, cloud infrastructure knowledge, and strong data modeling skills.