Skip to content

Staff Platform Engineer (Pacific Time Zone)

Lead technical direction for Komodo's core control plane (KMC/PSS, identity, subscriptions) and App Builder/Connector. Architect platform primitives, APIs, and AI tooling in a multi-tenant SaaS environment.

About the job

Responsibilities

  • Own and deliver major subscription service initiatives (e.g., My Subscriptions enhancements, self-service SAML SSO, new admin workflows) to improve time-to-onboard and operator efficiency
  • Architect and deploy custom MCP (Model Context Protocol) tooling and autonomous agents for seamless data interoperability between internal systems and AI models
  • Contribute to development of secure, reliable API endpoints and platform services to make platform infrastructure highly reliable, easy to maintain, and cost-effective
  • Define and improve monitoring, alerting, and observability for platform services; participate in on-call rotation
  • Act as escalation point for complex, cross-system incidents touching KMC, Connector, and App Builder
  • Collaborate with solution architects, forward-deployed engineers, customer success, and product teams to support new feature rollouts and integrate authentication services
  • Implement measures to enhance developer experience including well-documented code, clear API documentation, efficient debugging tools, and adoption of AI-powered development workflows (Cursor, Claude, Gemini)

Requirements

  • Expert-level backend engineering experience (Python + FastAPI preferred) building and operating APIs and microservices at scale with strong debugging and technical troubleshooting skills
  • Proven track record designing and evolving core platform primitives (authentication/authorization, orgs/accounts, subscriptions, or similar control-plane systems) in a multi-tenant SaaS environment
  • Profound experience with AWS core services and ability to architect secure, scalable, and cost-efficient solutions
  • Familiarity with modern data platforms and workflow tools (Snowflake, Airflow, Spark) and understanding of observability (metrics, logs, traces, events) across complex distributed systems
  • Ability to design solutions balancing performance, cost, maintainability, reliability, and developer experience
  • Demonstrated ability to drive complex, cross-team projects, mentor engineers, and collaborate with product managers, data scientists, and customer-facing teams
  • Comfort leveraging AI tools (Cursor, ChatGPT, Gemini) and interest in integrating LLM-based workflows into the SDLC

Nice-to-Haves

  • AWS cloud infrastructure certification
  • Experience with data privacy concerns such as HIPAA or GDPR
  • Working knowledge of data modeling and storage across relational (PostgreSQL), NoSQL (DynamoDB, Redis), and MPP databases (Snowflake, Redshift)
  • Prior experience in platform or infrastructure teams supporting internal and external developers, including SDKs, CLIs, and self-service tooling

Skills

Python, FastAPI, AWS, Snowflake, Airflow, Spark, Postgres, DynamoDB, Redis, HIPAA, GDPR

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Komodo Health

Komodo Health

United States

Staff Infrastructure Engineer
$187k+/yrRemote8+ YOEDevOps / SRE

Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.

VGS

VGS

United States
Senior Staff Infrastructure Engineer
$185k+/yrRemote10+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.