Software Engineer, Infrastructure
Infrastructure Software Engineer building scalable backend systems, observability, and developer tools on the Foundation team. Lead projects on database sharding, event bus, search scaling, and latency; mentor engineers. Requires 5+ years with distributed systems, NoSQL (MongoDB), IaC (Terraform), and architecture ownership.
About the job
Responsibilities
- Serve as a Lead and senior member of the team, leading initiatives and key projects.
- Provide guidance and mentorship to team members, fostering a culture of collaboration and technical excellence.
- Analyze, design, develop, maintain and improve software infrastructure and platform to expand its capabilities.
- Measure, report and drive improvements on scalability, performance, and availability.
- Lead complex projects and efforts to improve our ability to respond to impaired production systems.
- Lead, plan, and execute large scale system changes to meet business objectives.
- Tackle ambiguous engineering problems by designing well architected solutions.
- Participate in cross-team initiatives to drive engineering best-practices.
- Collaborate with the CTO and other technical leaders to set the vision for strategic development efforts.
- Work with various engineering teams to bridge gaps and lead initiatives across backend systems and our data platform.
- Proactively identify systemic inefficiencies, design and implement improvements.
- Become a trusted coach and mentor, actively building strong technical leaders.
- Collaborate with the InfoSec team to drive compliance, observability and automation for the security of our platform.
- Work closely with the Security team to address gaps and vulnerabilities during audits to satisfy compliance requirements.
- Manage and coordinate security vulnerabilities and upgrade schedules for EOL (End of Life) software.
- Lead or assist in security investigations as needed.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
- 5+ years experience building and managing large scale, highly available, distributed web applications.
- Mastery of a high-level programming language like Go, Python, JavaScript, etc.
- Experience with NoSQL databases like MongoDB.
- Own technical architecture discussions and leading decisions for an engineering organization.
- Lead and mentor engineers through technical projects and initiatives.
- Proficiency with Infrastructure as Code (IaC) tools, including Terraform.
- Experience designing sharding configurations for databases.
- Experience developing internal tools for others.
- Experience creating SLAs, SLOs, SLIs.
- HIPAA Compliance: All roles at Kustomer may involve handling sensitive personal data.
Nice-to-Haves
- Experience with our tech stack: Javascript (React/node.js), Go, Python, AWS Cloud, MongoDB, Redis, Elasticsearch, Terraform, Kafka, Kinesis.
Compensation and Benefits
- Competitive salaries and stock options.
- In the U.S.: 100% healthcare coverage, 401K, WiFi and Mobile reimbursement, and a generous vacation policy.
Skills
Go, Python, JavaScript, MongoDB, Terraform, AWS, Elasticsearch, Kafka, Redis, Kinesis
Similar jobs
DevOps / SRE jobsBuilds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including production services, data pipelines, observability, security, and CI/CD. Requires production LLM experience, strong AWS expertise, Python data-pipeline skills, and familiarity with RAG and LangChain.
Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.