Staff Software Engineer – AI Enablement
Leads the architecture and hands-on development of internal developer platforms, AI-agent workflows, and distributed automation infrastructure. Requires 8+ years of software engineering experience, Staff-level technical leadership, and strong backend, microservices, and event-driven architecture expertise.
About the job
Responsibilities
- Build internal developer platforms, service scaffolding, and automated CI/CD tooling to accelerate software delivery.
- Design multi-agent automation systems and adaptive conversational flows using LangGraph and foundation model APIs.
- Create domain-specific languages (DSLs) and finite-state-machine (FSM) abstractions for complex multi-agent platform tasks.
- Implement near-real-time memory (NRTM) and shared session management for persistent state across developer tools, bots, and internal APIs.
- Build resilient distributed systems using Redis, Kafka, and custom dead-letter-queue (DLQ) bridges.
- Establish AI guardrails, prompt-engineering practices, logging, and safety standards to protect sensitive platform data.
- Instrument microservices telemetry using sidecar patterns, data streams, and automated health checks for Datadog dashboards.
- Mentor engineers, lead architecture reviews, and establish engineering standards for AI agents and platform systems.
- Define and monitor developer-experience metrics such as onboarding time, deployment success rates, and task automation velocity.
Requirements
- Bachelor’s or graduate degree in Computer Science, Software Engineering, or a related technical field.
- 8+ years of professional software engineering experience developing backend systems, distributed architectures, or core platform tools.
- 2+ years operating at a Staff or Principal engineering scope.
- Formal experience mentoring engineers, leading architecture reviews, or serving as the technical lead for a multi-engineer project team.
- Strong hands-on coding skills in Java, Python, or Go.
- Deep experience with microservices architecture, RESTful web services, and event-driven architectures.
Nice-to-haves
- Experience building agentic workflows, multi-agent systems, and LLM integrations with LangGraph, AWS Bedrock, OpenAI API, or Anthropic API.
- Extensive practical experience with Terraform for cloud provisioning, state management, and scalable infrastructure deployment.
- Experience developing DSLs, FSMs, and dynamic agent-routing algorithms.
- Experience with Redis, Kafka, PostgreSQL, NoSQL databases, and streaming platforms.
- Experience with Docker, Kubernetes, and AWS cloud deployments.
- Experience implementing dead-letter queues, sidecar monitoring, automated secret management, and self-healing mechanisms.
- Ability to mentor engineers, influence architectural decisions, and partner with product and security teams.
Compensation and benefits
- Annual base salary: $217565–$260000.
- Base salary excludes bonus, sales incentives, equity, and benefits; final offers may vary based on qualifications, experience, skills, education, training, geographic location, and role.
- Benefits include medical, dental, and vision coverage; health savings and flexible spending accounts; life and AD&D insurance; 401(k) with company match; parental leave; paid time off and company holidays; disability, accident, and critical illness insurance; referral bonuses; employee assistance; pet insurance; travel assistance; wellbeing and childcare discounts; benefit advocacy; and learning and development benefits.
Skills
Java, Python, Go, LangGraph, Aws Bedrock, OpenAI API, Claude API, Terraform, Redis, Kafka, Postgres, Docker, Kubernetes, AWS, Datadog
Similar jobs
Backend Engineering jobsStaff Software Engineer building distributed backend platforms and infrastructure for customer service and compliance workflows. The role requires 8+ years of experience, deep Go expertise, cloud and DevOps skills, architectural leadership, and engineer mentorship.
Own and evolve Coinbase’s large-scale application routing infrastructure across Kubernetes, Envoy, and Istio. The role requires 8+ years in software and infrastructure engineering, strong systems programming and AWS experience, and expertise in distributed systems and service networking.
Leads the architecture and development of production agentic AI systems, distributed infrastructure, APIs, orchestration, guardrails, and evaluation frameworks for scalable automation. Requires 8+ years of software engineering experience and demonstrated expertise shipping AI/ML applications and reliable LLM-based systems.
Designs and operates high-throughput, low-latency ad-serving infrastructure, including GPU-based model inference and feature stores. Requires 10+ years of industry experience, strong computer science fundamentals, and a master’s degree in computer science or equivalent experience.
Leads the architecture and development of Crusoe Cloud’s software-defined networking infrastructure, including Linux kernel, driver, and packet-processing systems. Requires 8+ years of high-performance networking experience and expertise in C/C++, Linux internals, kernel bypass, and network accelerators.