Skip to content
DockerDocker

Staff Software Engineer, Agentic Platform

As a Staff Software Engineer on the Agentic Platform team, you will build foundational infrastructure for AI-driven workflows, focusing on agent execution runtime, orchestration, and cloud infrastructure. This role involves high ownership, driving architectural decisions, and mentoring junior engineers.

About the job

Responsibilities/What you'll work on:

Agent Workflow & Orchestration

  • Design and operate the core agent execution runtime responsible for scheduling, state management, and lifecycle management of long-running agentic workflows
  • Build robust multi-agent coordination patterns: task handoff, agent memory (short-term and long-term), tool use, and workflow branching at scale
  • Develop context window management strategies and session persistence layers for stateful agent interactions
  • Build tooling for prompt engineering as a first-class engineering discipline — versioning, testing, and evaluation of prompts at scale
  • Build platform capabilities that support developers working in AI-assisted coding workflows, including IDE integrations, local-first development environments, and fast iteration loops

Cloud Infrastructure & Service Ownership

  • Own and operate Agentic Platform services in AWS or OCI infrastructure provisioning, scaling, cost management, and reliability
  • Provision and manage cloud infrastructure using Terraform; manage Kubernetes application packaging and deployment with Helm
  • Participate in the 24/7 on-call rotation This role may require participation in a 24/7 on-call rotation for the Agentic Platform; carry genuine pager responsibility for the services you build and operate
  • Define and uphold SLOs; lead incident response, blameless post-mortems, and drive continuous reliability improvements
  • Instrument systems for observability: distributed tracing, structured logging, metrics dashboards, and alerting

Technical Leadership

  • As a Staff Engineer, partner with engineering leadership to set technical direction and serve as a guide and mentor as the team grows
  • Drive architectural decisions that balance velocity with long-term maintainability across a distributed, cloud-native stack
  • Collaborate cross-functionally with product managers, designers, and partner engineering teams to integrate agentic capabilities into the broader developer platform
  • Contribute to a culture of engineering excellence through design reviews, RFC processes, and mentorship

Qualifications for this role

Required:

  • 8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering.
  • Cloud Platform Expertise (AWS/OCI/Azure/GCP): Proven, hands-on experience operating production services in AWS or Oracle Cloud Infrastructure compute, networking, managed services, IAM, and cost management. This is a must-have; the Agentic Platform is a cloud-native service running 24/7.
  • Service Ownership in a Cloud Setting: You have owned production services end-to-end — on-call, incident response, SLO definition, and post-mortems. You don't just build; you run what you build.
  • Distributed Systems Design: Deep understanding of fault tolerance, consistency, observability, and scalability in cloud-native environments
  • Backend Engineering Proficiency: Strong proficiency in at least one backend language used for systems work — Go, Python, Rust, or Java
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience

Strongly Preferred:

  • Go: Professional proficiency in Go — Docker's primary language for backend systems
  • Infrastructure as Code: Experience with Terraform for cloud infrastructure provisioning and Helm for Kubernetes application packaging and deployment
  • Data Infrastructure: Experience with PostgreSQL and Redis / Pub-Sub patterns for state management, caching, and event-driven agent workflows
  • MCP & Agent Tooling: Experience with MCP (Model Context Protocol) server design and integration
  • Container & Orchestration: Docker, Kubernetes, or equivalent — especially in the context of agent sandboxing and secure code execution environments
  • AI-assisted development tools: Familiarity with Cursor, Claude Code, Copilot, Windsurf, etc. and the developer personas using them
  • Agent Evaluation: Experience with LLM-as-judge frameworks, behavioral regression testing, and golden dataset management
  • Agent Systems Experience: Hands-on experience building or operating AI agent systems — including multi-agent orchestration, tool use, memory systems, or agent evaluation frameworks
  • Open Source: Contributions or community engagement on relevant open source projects

Perks

  • Freedom & flexibility; fit your work around your life
  • Designated quarterly Whaleness Days plus end of year Whaleness break
  • Home office setup; we want you comfortable while you work
  • 16 weeks of paid Parental leave (after 6 months of employment)
  • Technology stipend equivalent to $100 USD net/month
  • PTO plan that encourages you to take time to do the things you enjoy
  • Equity; we are a growing start-up and want all employees to have a share in the success of the company
  • Docker Swag
  • Medical benefits, retirement and holidays vary by country
  • Remote-first culture, with offices in Seattle and Paris

Skills

AWS, Oci, Go, Python, Rust, Java, Terraform, Helm, Postgres, Redis, Docker, Kubernetes

Censys

Censys

United States

Staff Scanning Engineer
$172k+/yrRemote10+ YOEBackend Engineering

Designs and operates Internet-scale scanning, DNS, attribution, and data pipelines, with deep ownership of distributed backend systems and production reliability. Requires 10+ years of software engineering experience, strong Go expertise, cloud and streaming infrastructure knowledge, and the ability to mentor engineers.

Grafana Labs

Grafana Labs

United States

Staff Backend Engineer - Grafana App Platform| US| Remote
$175k+/yrRemote7+ YOEBackend Engineering

Staff backend engineer responsible for evolving Grafana into a scalable, multi-tenant observability application platform. The role requires production operations experience, distributed-systems expertise, strong communication, and familiarity with Go or willingness to learn it.

Turion Space

Turion Space

Irvine, CA

Staff Ground Software Engineer
$175k+/yrOn-site8+ YOEBackend Engineering

Leads architecture and development of mission-critical backend systems for spacecraft command, telemetry, mission planning, and operations. Requires 8+ years of software development experience, distributed-systems and cloud-native expertise, and technical leadership.

Pinterest

Pinterest

United States

Staff Software Engineer, TwoTwenty
$177k+/yrRemote7+ YOEBackend Engineering

Leads the technical direction and hands-on development of scalable backend platforms for Pinterest’s AI-driven products. Requires extensive backend and distributed-systems experience, strong Python skills, production LLM or ML experience, and cross-functional technical leadership.

Pinterest

Pinterest

United States

Staff Software Engineer, Storage Services
$177k+/yrRemote8+ YOEBackend Engineering

Leads the technical strategy, architecture, and operation of Pinterest’s large-scale storage infrastructure supporting SQL and graph workloads. Requires 8+ years of backend distributed-systems experience, technical leadership, and expertise in production reliability, performance tuning, and modern database technologies.