Senior Software Engineer, Observability
Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.
About the job
Responsibilities
- Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows.
- Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services.
- Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm.
- Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis.
- Define observability best practices.
Requirements
- Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure).
- Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm).
- Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying.
- Deep understanding of containerization (Docker) and orchestration (Kubernetes).
- Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows.
- Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data.
Preferred
- Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines.
- Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing.
- Contributions to open-source observability projects.
- Familiarity with security monitoring and compliance frameworks.
Compensation
- US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits.
Skills
Prometheus, Grafana, ClickHouse, OpenTelemetry, Go, Python, Terraform, Ansible, Helm, Kubernetes, Docker, AWS, GCP, Azure
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.