Staff Software Engineer, AI Developer Tools
Staff-level engineer architecting AI-native developer tools and infrastructure to accelerate engineering velocity across Gusto. Requires 8+ years experience building production AI systems with deep expertise in LLMs, RAG, and multi-agent workflows.
About the job
Responsibilities
- Architect AI Core Infrastructure: Design, scale, and maintain the internal platform and APIs that integrate LLMs into the developer workflow, with a heavy focus on autonomous agents, custom retrieval-augmented generation (RAG), and smart context-delivery systems.
- Drive Technical Leadership & Vision: Establish engineering best practices for building with AI at Gusto, including robust evaluation frameworks (evals), guardrails for code safety, and strategies for optimizing model latency and token spend.
- Collaborate & Evangelize: Partner closely with product engineering teams to uncover systemic friction points, and collaborate with core Infrastructure and Security teams to ensure all AI tooling is scalable, compliant, and secure.
- Optimize Feedback Loops: Identify and eliminate critical bottlenecks in the local-to-production lifecycle at scale across Gusto.
- Maintain Enterprise Reliability: Treat AI infrastructure with production-grade rigor. Monitor system health, diagnose complex integration or networking issues, and mitigate model drift or downtime.
Requirements
- 8+ Years of Systems Expertise: Seasoned engineer comfortable navigating massive, complex codebases and delivering production-grade, distributed software.
- Production AI Experience: Proven track record of moving AI beyond prototypes. Deep understanding of prompt engineering, fine-tuning, RAG architectures, and orchestrating multi-agent workflows.
- Developer-First Mindset: Passionate about Developer Experience and Productivity. Understand what makes an internal tool frictionless, fast, and empowering for other engineers.
- Strategic Execution: Think deeply about data privacy, security, cost management, and systemic reliability when deploying AI solutions.
- Exceptional Communication: Translate bleeding-edge AI capabilities into clear, actionable infrastructure strategies, mentoring junior engineers and aligning cross-functional stakeholders.
- Working as a Team: Work across team boundaries and bring other teams and groups along towards strategic vision.
Compensation & Benefits
- Cash compensation targeted at $180,000/yr to $200,000/yr in Denver & most remote locations, $220,000/yr to $245,000/yr for San Francisco, New York & Seattle.
- Stock equity is additional.
- All full-time employees receive competitive base pay, benefits, and equity (RSUs).
Skills
LLMs, RAG, Prompt Engineering, Fine-Tuning, Multi-Agent Systems, AI Infrastructure, Distributed Systems, Evaluation Frameworks, CI/CD, Developer Productivity Tools
Similar jobs
DevOps / SRE jobsBuild and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.