Manager, Software Engineering - Observability
Lead a team building Figma's observability platform with Datadog and OpenTelemetry, driving instrumentation strategy, cost optimization, and AI-powered anomaly detection. Requires 4+ years managing infrastructure teams and deep distributed systems expertise.
About the job
What you’ll do at Figma:
- Lead and grow a team responsible for reliability, scalability, and evolution of observability and cost engineering platforms
- Own and operate core observability stack including Datadog, ensuring high availability and data quality
- Define technical strategy for instrumentation standards, libraries, agents, and operators
- Implement AI-driven approaches to anomaly detection, root cause analysis, and automation
- Establish frameworks for cost attribution, budgeting, forecasting, and alerting
- Partner with infrastructure, product engineering, finance, and security teams
- Optimize observability footprint and spend
- Coach and mentor engineers on career development and technical leadership
We'd love to hear from you if you have:
- 4+ years leading infrastructure, observability, or platform engineering teams
- Deep experience with observability platforms (Datadog, OpenTelemetry)
- Strong understanding of distributed systems, instrumentation, SLOs, and incident response
- Experience driving cost transparency and accountability in cloud environments
- Ability to set technical direction and drive cross-functional alignment
Nice-to-haves:
- Experience with company-wide observability standards and integrations
- Background in cost optimization and vendor negotiations
- AI/ML for anomaly detection or automation
- Familiarity with OpenTelemetry across languages
- Scaling and mentoring high-performing teams
Skills
Datadog, OpenTelemetry, Distributed Systems, Instrumentation, Slo Design, Incident Response, Cost Attribution, Ai Anomaly Detection, Kubernetes, Cloud Cost Management
Similar jobs
Engineering Management jobsLeads and scales a customer-facing Applied AI Engineering team serving high-growth startups, guiding customers from experimentation to production while building repeatable deployment mechanisms. Requires technical depth in AI/ML platforms and experience leading teams in startup-focused, ambiguous environments.
Leads the Web Infrastructure team responsible for Notion’s web client architecture, performance, reliability, and shared design systems. The role manages senior engineers and managers, drives execution and planning, and partners across the organization on technical and organizational practices.
Leads multiple engineering teams, combining people management with hands-on technical leadership, architecture, and coding. The role requires at least four years of software engineering experience and a demonstrated ability to build high-performing teams in complex startup environments.
Leads technical strategy and a multidisciplinary engineering team responsible for billing, accounting, eligibility, and enrollment systems. The role requires 5+ years of engineering management experience, strong distributed-systems expertise with Python or Go, and experience developing engineering leaders.
Leads and builds a platform team responsible for infrastructure, developer productivity, and SRE/production support. The role combines hands-on technical work with team building and requires cloud, infrastructure troubleshooting, large-scale operations, and SRE experience.