TLM, Production Engineering
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
About the job
Responsibilities
- Build and operate critical infrastructure across compute, storage, messaging, and observability systems handling financial transactions at scale.
- Drive architectural changes to improve reliability, scalability, and performance; own solutions through implementation and resolution.
- Partner with product teams during design reviews and establish reusable golden paths.
- Contribute to Ramp’s cellular architecture, supporting international markets, regulated environments, and enterprise-grade SLAs.
- Develop infrastructure patterns for AI-powered products and reusable platform foundations.
- Build developer tooling and self-service infrastructure for cost, performance, and reliability visibility.
- Participate in an on-call rotation and eliminate incident root causes.
- Lead cross-team architectural reviews and company-wide reliability and scalability initiatives.
Requirements
- 3+ years of software engineering experience shipping high-quality architectures for critical systems.
- Strong software engineering fundamentals, including clean, well-tested, production-ready code.
- Hands-on experience with distributed systems at production scale.
- Experience with at least one major cloud provider; AWS preferred.
- Familiarity with observability practices, including SLOs, error budgets, alerting, and dashboards.
- Experience leading technical projects end to end, including cross-team coordination.
- Comfort using AI tooling and coding agents in everyday engineering workflows.
Nice-to-haves
- Experience with cellular or multi-tenant architecture patterns.
- Experience with workflow orchestration systems such as Temporal.
- Contributions to developer experience or internal platform tooling.
- Experience in fintech, payments, or regulated industries, including FedRAMP or SOC 2.
Compensation and Benefits
- Salary range: $168,000–$324,500.
- Flexible paid time off.
- Centralized home-office equipment ordering.
- Health and wellness stipend.
- Budget for intra-office travel.
- Weekly benefits and employee programs are available to full-time employees.
Skills
Distributed Systems, AWS, Cloud Computing, Observability, SLOs, Error Budgets, Terraform, Temporal, Cellular Architecture, Multi-Tenant Architecture, AI Tools, CI/CD, Container Orchestration, Networking, Load Balancing
Similar jobs
DevOps / SRE jobsLeads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.