Staff Infrastructure Engineer
Staff Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents that support 911 emergency call centers. Requires 6+ years in infrastructure/platform/backend roles with experience in reliability and scale.
About the job
Responsibilities
- Build out analytics infrastructure using ClickHouse to handle massive increases in traffic and increasingly complex queries.
- Create great instrumentation, logging, and monitoring tools to give as much visibility as possible into where calls go wrong in production.
- Develop testing frameworks, CLIs, automations, etc. that help the team move quickly with AI coding tools without sacrificing reliability or confidence.
- Build the systems, tools, and abstractions that make other engineers more productive and the platform more reliable.
Requirements
- 6+ years of experience in infrastructure, platform, or backend engineering roles, ideally at companies where reliability and scale are serious challenges.
- Comfortable across areas of the stack from backend application code to cloud infrastructure.
- Thrive on autonomy and constantly look for new areas to improve.
- Excited by the challenges of building AI systems at scale.
Nice-to-Haves
- Experience with ClickHouse.
- Experience with AI systems at scale.
Compensation and Benefits
- Base salary: $200,000 - $300,000 (total compensation may include equity).
- Comprehensive Medical, Dental, Vision & Life insurance.
- 401(k).
- Unlimited PTO.
- Company-wide offsites.
- Equipment stipend.
- Relocation assistance.
- Daily delivered lunches.
- Office in Seattle.
- Start-up Equity.
Skills
ClickHouse, Infrastructure, Observability, Logging, Monitoring, Testing Frameworks, Cli, Automation, Backend Engineering, Cloud Infrastructure, Ai Systems
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.