Senior Software Engineer, Release Infra
Senior Software Engineer building and operating Brex's release infrastructure, CI/CD pipelines, observability, and incident management systems. Requires 7+ years experience with backend languages, Kubernetes, cloud platforms, and SRE practices to ensure safe, scalable deployments.
About the job
Responsibilities
- Design, build, and maintain the release infrastructure that powers Brex’s deployment pipelines and incident workflows
- Drive technical strategy and architecture for release and observability systems, making them more scalable, reliable, and secure
- Collaborate with product, engineering, and operations partners to ensure Brex’s releases are safe, predictable, and low-friction
- Identify and deliver improvements to the end-to-end release process (from code merge to production) to reduce risk and cycle time
- Build and evolve tooling for observability and incident response, enabling fast detection, triage, and resolution
- Proactively identify and mitigate risks in our release and infrastructure stack, including performance, reliability, and security concerns
- Define, instrument, and monitor key metrics for release engineering (e.g., deployment frequency, change failure rate, MTTR) and use them to guide improvements
- Partner with other infrastructure and product teams to debug complex production issues and drive long-term fixes
- Contribute to and champion best practices in release engineering, reliability, and operational excellence across the organization
- Mentor other engineers on the team, providing technical guidance and code reviews to elevate the overall quality of our infrastructure
- Stay up-to-date on emerging tools and practices in release engineering, observability, and SRE, and bring relevant ideas into Brex’s stack
Requirements
- 7+ years of professional experience designing, building, and operating backend or infrastructure systems in production
- Strong proficiency in backend programming languages (e.g., Go, Java, Kotlin, or Python) with a focus on reliability and performance
- Hands-on experience with CI/CD and release pipelines (e.g., GitHub Actions, CircleCI, Buildkite, Argo, Spinnaker, Jenkins) including build, test, and deployment automation
- Experience architecting and operating scalable, high-availability distributed systems on cloud platforms (e.g., AWS, GCP, Azure)
- Deep familiarity with containerization and orchestration (e.g., Docker, Kubernetes) and infrastructure-as-code (e.g., Terraform, CloudFormation)
- Experience designing and maintaining observability tooling (metrics, logs, tracing) and integrating it into incident response workflows
- Strong understanding of reliability and SRE practices, including SLIs/SLOs, error budgets, and incident management best practices
- Experience designing and optimizing data storage systems (SQL and/or NoSQL) for operational and observability use cases
- Proven track record of improving release processes (e.g., reducing deployment risk, increasing deployment frequency, automating rollbacks)
- Comfort working cross-functionally with product and other engineering teams to debug complex production issues and ship changes safely
- Strong communication and collaboration skills, including writing clear design docs and driving technical decisions across teams
Nice-to-Haves
- Experience with emerging tools and practices in release engineering, observability, and SRE
Compensation
The expected salary range for this role is $192,000 - $240,000. However, the starting base pay will depend on a number of factors, including the candidate’s location, skills, experience, market demands, and internal pay parity. Depending on the position offered, equity and other forms of compensation may be provided as part of a total compensation package.
Skills
Go, Java, Kotlin, Python, GitHub Actions, CircleCI, Buildkite, Argo, Spinnaker, Jenkins, AWS, GCP, Azure, Docker, Kubernetes
Similar jobs
DevOps / SRE jobsOwn the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.
Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.