Skip to content
NumeralNumeral

Software Engineer (Infra)

Builds and scales core infrastructure for a high-growth AI tax platform, focusing on reliable APIs, data pipelines, observability, and fault-tolerant distributed systems. Requires 7+ years experience with Node.js, PostgreSQL, Redis, AWS, Kubernetes, and observability tools.

About the job

Responsibilities

  • Design and build highly scalable, secure, and reliable infrastructure to support critical APIs, services, and data pipelines.
  • Lead infrastructure architecture decisions for performance, observability, and fault tolerance.
  • Drive improvements in system reliability, monitoring, and incident response.
  • Collaborate closely with product, platform, and data engineering teams to enable rapid feature development without compromising stability.
  • Automate infrastructure management through infrastructure-as-code, CI/CD pipelines, and container orchestration.
  • Champion best practices in operational excellence, disaster recovery, and distributed systems design.

Requirements

  • 7+ years of experience building and scaling backend or infrastructure systems in high-growth environments.
  • Proficiency in Node.js, PostgreSQL, Redis, and AWS (or equivalent cloud platforms).
  • Experience managing containerized workloads using tools like Kubernetes, ECS, or similar.
  • Expertise in observability tooling (e.g., Datadog, Prometheus, OpenTelemetry) and monitoring pipelines.
  • Proven ability to design fault-tolerant, distributed systems at scale.
  • Strong debugging and performance optimization skills across the stack.
  • Exceptional collaboration and communication skills, with experience partnering cross-functionally.
  • Startup mindset: Not scared of ambiguity and hungry for rapid growth.
  • Intensity & Ownership: This is not a 9-5 — we’re scaling rapidly and have a massive opportunity ahead.
  • Customer Obsession: You deeply care about the user experience and solving their problems.

Nice-to-Haves

  • Experience in payments, tax, accounting, or regulatory tech.
  • Infrastructure or platform engineering background (e.g., CI/CD, observability).
  • Experience scaling monolith-to-service architectures or event-driven systems.
  • Familiarity with GraphQL, Kafka, Terraform, or container orchestration tools.

Compensation & Benefits

  • Competitive salary and equity.
  • Full medical, dental, and vision coverage.
  • Wellness perks like Headspace and the Peloton One App.
  • 401(k).
  • Lunch and snacks when you’re in the office.
  • Regular team offsites and company events.

Skills

Node.js, Postgres, Redis, AWS, Kubernetes, ECS, Datadog, Prometheus, OpenTelemetry, Terraform, GraphQL, Kafka, CI/CD

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Onebrief

Onebrief

Colorado Springs, CO

Senior Site Reliability Engineer, Colorado Springs
$180k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.