Senior Software Engineer - Observability
Develops large-scale observability infrastructure for Databricks products and platform systems, including logging, metrics, tracing, dashboards, and alerting. The role requires 5+ years of production software engineering experience and familiarity with distributed systems and observability practices.
About the job
Responsibilities
- Establish standards for logging, metrics, and tracing.
- Collaborate with different teams to identify metrics that show how systems and their subcomponents are performing.
- Build tooling and infrastructure that enable components to efficiently emit, aggregate, and store metrics for dashboards and alerting.
- Contribute to and execute the technical roadmap to ensure scalability, performance, and reliability.
- Participate in on-call rotations and reduce incident response times.
- Optimize platform and infrastructure by analyzing system expenses, improving visibility, promoting mindful usage, enforcing retention policies, streamlining queries, and right-sizing resources to reduce observability costs.
Requirements
- Bachelor's degree or higher in Computer Science or a related field.
- 5+ years of production-level experience with Python, Java, Scala, C++, or similar languages.
- Experience with software development in large-scale distributed systems.
- Familiarity with metrics collection, health monitoring, and observability tools.
Benefits
- Comprehensive benefits and perks tailored to employees' regional needs.
Skills
Python, Java, Scala, C++, Distributed Systems, Metrics Collection, Health Monitoring, Observability Tools, Logging, Tracing, Alerting, Dashboards, On-Call Operations
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.