Software Engineer, Core Infrastructure
Builds and scales core cloud infrastructure for high availability, performance, and developer productivity. Collaborates across teams to evolve backend architecture, automate workflows, and maintain production systems using Kubernetes, Docker, Postgres, and Node.
About the job
Responsibilities
- Scale Retool’s core cloud platform for high availability and performance globally
- Work collaboratively with the rest of the engineering team to deliver infrastructure for core and emerging products
- Evolve our backend architecture/infrastructure for both cloud and on-premise deployments
- Work with the team to set and prioritize our roadmap to maximize customer impact
- Define and automate developer process/workflow
- Support our systems in production
- Develop new data solutions
- Build monitoring and observability for production systems
Requirements
- Track record of delivering engineering projects and process improvements
- Experience scaling cloud infrastructure
- Enjoy building and productionizing developer productivity tools, frameworks, and other aspects of platform engineering
- Experience with inner workings of Linux, containers (Docker, containerd), and container orchestration technologies (e.g. docker-compose, Kubernetes)
- Track record of building productive, collaborative relationships, both within an engineering org and across the broader company
- Enjoy the ambiguity and high-ownership culture of early-stage startups
- Pragmatic, solution-oriented, and scrappy
- Enjoy working collaboratively with a broad range of job functions and roles
- Experience with our tech stack: Node, Postgres, Docker, Kubernetes
- Experience scaling relational databases (Postgres preferred)
- Good knowledge of cloud, on-prem, traffic routing, service architecture in multi-region setup
Skills
Kubernetes, Docker, Postgres, Node.js, Linux, Cloud Infrastructure, Containerd, Docker-Compose, Relational Databases
Similar jobs
DevOps / SRE jobsInfrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.