Staff Platform Engineer (Pacific Time Zone)
Lead technical direction for Komodo's core control plane (KMC/PSS, identity, subscriptions) and App Builder/Connector. Architect platform primitives, APIs, and AI tooling in a multi-tenant SaaS environment.
About the job
Responsibilities
- Own and deliver major subscription service initiatives (e.g., My Subscriptions enhancements, self-service SAML SSO, new admin workflows) to improve time-to-onboard and operator efficiency
- Architect and deploy custom MCP (Model Context Protocol) tooling and autonomous agents for seamless data interoperability between internal systems and AI models
- Contribute to development of secure, reliable API endpoints and platform services to make platform infrastructure highly reliable, easy to maintain, and cost-effective
- Define and improve monitoring, alerting, and observability for platform services; participate in on-call rotation
- Act as escalation point for complex, cross-system incidents touching KMC, Connector, and App Builder
- Collaborate with solution architects, forward-deployed engineers, customer success, and product teams to support new feature rollouts and integrate authentication services
- Implement measures to enhance developer experience including well-documented code, clear API documentation, efficient debugging tools, and adoption of AI-powered development workflows (Cursor, Claude, Gemini)
Requirements
- Expert-level backend engineering experience (Python + FastAPI preferred) building and operating APIs and microservices at scale with strong debugging and technical troubleshooting skills
- Proven track record designing and evolving core platform primitives (authentication/authorization, orgs/accounts, subscriptions, or similar control-plane systems) in a multi-tenant SaaS environment
- Profound experience with AWS core services and ability to architect secure, scalable, and cost-efficient solutions
- Familiarity with modern data platforms and workflow tools (Snowflake, Airflow, Spark) and understanding of observability (metrics, logs, traces, events) across complex distributed systems
- Ability to design solutions balancing performance, cost, maintainability, reliability, and developer experience
- Demonstrated ability to drive complex, cross-team projects, mentor engineers, and collaborate with product managers, data scientists, and customer-facing teams
- Comfort leveraging AI tools (Cursor, ChatGPT, Gemini) and interest in integrating LLM-based workflows into the SDLC
Nice-to-Haves
- AWS cloud infrastructure certification
- Experience with data privacy concerns such as HIPAA or GDPR
- Working knowledge of data modeling and storage across relational (PostgreSQL), NoSQL (DynamoDB, Redis), and MPP databases (Snowflake, Redshift)
- Prior experience in platform or infrastructure teams supporting internal and external developers, including SDKs, CLIs, and self-service tooling
Skills
Python, FastAPI, AWS, Snowflake, Airflow, Spark, Postgres, DynamoDB, Redis, HIPAA, GDPR
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.