# Engineering Manager, Telemetry Agent and Edge

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** San Francisco, CA
**Role:** Engineering Management
**Salary:** $215k – $260k/yr
**Experience:** 7+ years
**Skills:** Kubernetes, Helm, Go, Rust, C++, OpenTelemetry, ebpf, telemetry, Observability, GPU
**Posted:** 2026-07-20

> Lead a team of 4-6 engineers building and operating Crusoe's telemetry agent for metrics/logs from hosts and GPUs. Own delivery of next-gen agent releases against fixed deadlines while hiring, coaching staff-level engineers, and maintaining high operational standards for fleet-wide software.

## Job Description

## What You'll Be Working On
- Grow and develop your team. Manage 4 to 6 engineers directly: 1:1s, career growth, performance, and team health, with the team growing over time.
- Own delivery. Plan, sequence, and de-risk a multi-phase release with a fixed finish line, and be accountable for shipping it.
- Set the technical direction for your area. Partner closely with the tech lead so they can focus on engineering rather than absorbing delivery and handoff work alone. Ask the hard questions and make sound tradeoff calls alongside your engineers.
- Keep the operational bar high. The agent runs on every customer node, so rollout safety, low overhead, and reliability are core to the job, not afterthoughts.
- Own the team's oncall rotation and the operational health of the agent fleet.
- Coordinate across team boundaries. Much of this team's work ships inside other teams' releases and hardware launch windows, so sequencing dependencies and escalating early is a core part of the job.
- Shape the team over time. Work with recruiting on sourcing, run a high-quality interview loop, close strong candidates, and onboard them well.
- Collaborate across functions. Partner with product, neighboring infrastructure teams, peer managers, and leadership to keep priorities aligned as the observability product expands.

## What You'll Bring to the Team
- You care about people. You want the engineers around you to grow, you give feedback that's both direct and kind, and you measure your own success through your team's.
- Technical depth. 5+ years of hands-on engineering experience, ideally in systems software: host agents, telemetry collectors, daemons, or other performance-sensitive software (Go, Rust, or C++ environments), and hands-on familiarity with Kubernetes and helm-based deployment.
- Experience managing engineers. 2+ years managing software engineers directly.
- Experience coaching senior and staff-level engineers, and comfort partnering with a strong tech lead rather than competing with them.
- A track record of shipping. You've delivered multi-phase projects against fixed deadlines, and you make honest calls early when a plan is at risk.
- Strong communication and judgment. You can align people across functions, explain tradeoffs to technical and non-technical partners alike, and give your team clarity about what matters and why.

## Bonus Points
- Experience building or operating host-level agents at fleet scale: safe rollouts and canarying, version management, self-update, crash recovery, and keeping resource overhead low on customer machines.
- Experience with hardware platform bring-up or new product introduction, validating software readiness on new GPU or server platforms.
- Familiarity with agent internals: collection scheduling, buffering and backpressure, local spooling, and graceful behavior when the control plane or network is unavailable.
- Data plane experience in a cloud provider or similar environment: hypervisors and virtualization, host networking, storage services, or other software in the customer serving path.
- Observability domain experience: OpenTelemetry, metrics and log pipelines, time series storage.
- Experience operating software in GPU or accelerated computing environments.
- Familiarity with kernel-level telemetry (eBPF, perf, tracing).
- Experience managing or scaling a team through growth.

## Benefits
- Competitive compensation and equity packages, Restricted Stock Units.
- Paid time off, paid holidays & leave of absence programs.
- Comprehensive health, dental & vision insurance, Employer contributions to HSA account.
- Paid parental leave, Paid life insurance, short-term and long-term disability.
- Professional development & tuition reimbursement.
- Mental health & wellness support.
- Commuter benefits (parking & transit), Cell phone stipend.
- 401(k) Retirement plan with company match up to 4% of salary.
- Volunteer time off.
- Global travel insurance & emergency assistance.
- Daily meals allowance.
- Additional perks & programs specific to location.

**Compensation Range:** up to $215,000 - $260,000 + Bonus. Restricted Stock Units are included in all offers.

## Similar roles

- [Software Engineering Manager, Database](https://hotfix.jobs/jobs/79fff42f-4f6f-43fa-8374-f92796c17a67) - LangChain - San Francisco, CA - $215k – $260k/yr
- [Senior Software Engineer, Test Infrastructure](https://hotfix.jobs/jobs/43a325b9-43ee-4084-9b48-d9ebe730918c) - Zoox - Foster City, CA - $215k – $237k/yr
- [Senior Engineering Manager, AI Product](https://hotfix.jobs/jobs/b5a1b7ec-b04b-4b3b-984f-14e754acf32f) - Mozilla - Remote - $215k – $240k/yr
- [Senior Engineering Manager, AI Product](https://hotfix.jobs/jobs/ded4bce0-de96-4945-a53d-ad96c74fdb18) - Mozilla - Remote - $215k – $240k/yr
- [Manager, Software Engineering, Search Discovery](https://hotfix.jobs/jobs/ab512cf9-08fd-4364-81cd-83b41a3fcfbd) - Cribl - Remote - $215k – $255k/yr

**Apply:** https://hotfix.jobs/jobs/2250c5fe-bc6f-4a10-aadb-345ec7df1427
**Canonical:** https://hotfix.jobs/jobs/2250c5fe-bc6f-4a10-aadb-345ec7df1427