Software Engineer, AI Runtime & Platform Services
Build and maintain CrewAI's Python enterprise runtime layer (FastAPI, Celery, Redis) that turns open-source agents into secure, observable production systems. Requires strong backend/platform engineering experience with distributed systems, auth, and observability.
About the job
What You'll Do
- Build and maintain the Python enterprise runtime around CrewAI: FastAPI services, Celery workers, Redis-backed state, execution APIs, and deployment-facing tools.
- Extend open-source CrewAI behavior for enterprise environments while preserving compatibility with upstream framework changes.
- Own production execution flows: crew and flow kickoff, status, retries, cancellation, checkpoint restore and fork, chat/session state, and human-in-the-loop resume paths.
- Build secure integration surfaces: JWT auth, signed webhooks, token refresh, file handling, secret fetching, and workload identity across AWS, GCP, and Azure.
- Improve observability across distributed execution: OpenTelemetry traces, structured logs, Sentry, event tracking, and debuggability across API, worker, and platform boundaries.
- Maintain strong test coverage for async/runtime behavior using pytest, mypy, ruff, mocks/fakes, and e2e deployment harnesses.
- Partner with the Agent Management Platform team on API contracts, versioning, enterprise client behavior, deployment status, and failure reporting.
Requirements
- Strong Python backend/platform engineering experience, especially building production services rather than only libraries.
- Experience with FastAPI or similar API frameworks, Celery or other job systems, Redis, Pydantic, and typed Python.
- Good instincts for distributed systems: retries, idempotency, async execution, status tracking, race conditions, and failure recovery.
- Comfort with auth and security-sensitive systems: JWTs, webhooks, signatures, secrets, IAM/workload identity, and least-privilege thinking.
- Practical observability experience: tracing, structured logging, metrics, Sentry/OpenTelemetry, and debugging multi-service failures.
- Ability to work at the boundary between an open-source framework and a hosted enterprise platform without creating brittle coupling.
- Strong testing habits and comfort with CI, package/version management, and release discipline.
Nice-to-Haves
- Experience operating AI/agent runtimes, workflow engines, or distributed task systems.
- Cloud platform experience with AWS ECS/ECR, Kubernetes, Helm, GCP/Azure identity, or secret managers.
- Experience with enterprise SaaS constraints: auditability, tenant isolation, customer environments, deployment rollbacks, and supportability.
- Familiarity with Rails/SaaS platforms is useful, but not required.
Skills
Python, FastAPI, Celery, Redis, Pydantic, OpenTelemetry, Sentry, AWS, GCP, Azure, Kubernetes, Jwt, Pytest
Similar jobs
Backend Engineering jobsBuild and scale Ruby on Rails backend services for rewards, incentives, and loyalty features in a high-throughput platform. The role owns complex feature delivery, contributes to architecture, mentors junior engineers, and requires 3+ years of professional software engineering experience.
Build and operate backend capabilities for a cloud identity platform, including authentication flows, APIs, and developer experiences. The role requires 3+ years of experience with high-scale production systems and RESTful API development, with Go, TypeScript, and DynamoDB as preferred skills.
Build and scale reliable research infrastructure and distributed systems for evolving AI research workflows. The role independently leads complex technical projects, makes foundational architectural decisions, and partners with researchers and engineering teams.
Build and operate Internet-scale HTTP and TLS infrastructure, migrate services to a Rust-based proxy, and improve protocol performance. The role requires systems programming experience, strong reliability and security practices, and interest in open-source standards.
Build secure, scalable backend platforms and AI integrations connecting enterprise systems, tools, and data sources. The role requires 5+ years of backend engineering experience, distributed-systems expertise, cloud-native operations, authentication security, and technical leadership.