# Software Engineer - Observability

**Company:** [Baseten](https://hotfix.jobs/companies/baseten)
**Location:** San Francisco, CA, New York, NY
**Role:** Backend Engineering
**Salary:** $165k – $330k/yr
**Skills:** Python, Rust, Go, Prometheus, Grafana, ClickHouse, OpenTelemetry, telemetry pipelines, columnar storage, Distributed Systems, metrics, logging, tracing, slos, ai/llms
**Posted:** 2026-08-10

> Build foundational observability infrastructure, including high-throughput telemetry pipelines, storage, instrumentation, alerting, and SLO systems across multi-cloud environments. The role requires strong programming skills in Python, Rust, or Go and experience with observability platforms and large-scale data systems.

## Job Description

## Responsibilities
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across multi-cloud infrastructure.
- Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth.
- Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services.
- Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliability targets with minimal noise.
- Partner with Inference, Product, and Infrastructure teams to ensure observability solutions meet organizational needs.

## Requirements
- Deep experience with at least one observability signal area: metrics, logging, tracing, or error analytics; familiarity with the others.
- Understanding of high-throughput data pipelines, columnar storage engines, and the tradeoffs involved in ingesting and querying telemetry data at scale.
- Experience operating or building on observability platforms such as Prometheus, Grafana, ClickHouse, OpenTelemetry, or similar systems.
- Strong proficiency in at least one of Python, Rust, or Go.
- Excellent communication skills and interest in improving operational visibility and incident response.
- Interest in applying AI/LLMs to operational workflows such as automated root cause analysis, anomaly detection, or intelligent alerting.
- Excitement about building foundational infrastructure and working independently on ambiguous, high-impact technical challenges.

## Compensation and Benefits
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employees and dependents.
- Flexible PTO, including a company-wide winter break.
- Paid parental leave.
- Fertility and family-building stipend.
- Company-facilitated 401(k).

## Similar roles

- [Software Engineer](https://hotfix.jobs/jobs/aa5a26c4-e2c2-457c-bd37-9df0b3d44d3f) - Cardless - San Francisco, CA - $165k – $260k/yr
- [Software Engineer, Ads](https://hotfix.jobs/jobs/ee15cbc3-c47c-4f72-b218-5406d4b43a34) - Reddit - Remote - $164k – $230k/yr
- [Software Engineer, Networking](https://hotfix.jobs/jobs/4a079177-8578-4a2e-a09a-c2d2e6e0358e) - Tailscale - Remote - $163k – $226k/yr
- [Software Engineer, Secure Development Engineering](https://hotfix.jobs/jobs/07c544c5-19dd-4786-b94f-2ddb330d1cc3) - Airbnb - Remote - $162k – $186k/yr
- [Software Engineer, Dev API](https://hotfix.jobs/jobs/8ab669bc-c97a-4904-8c17-f15758eec248) - Ramp - New York, NY - $168k – $330k/yr

**Apply:** https://hotfix.jobs/jobs/87275682-f9eb-4541-aa59-caee6cf19e01
**Canonical:** https://hotfix.jobs/jobs/87275682-f9eb-4541-aa59-caee6cf19e01