Senior Platform Monitoring Engineer
Leads platform observability, monitoring, and incident response efforts, designing alerting and automation that improve reliability and customer experience. Requires 6+ years in SRE, DevOps, production engineering, or a similar role, along with cloud, container orchestration, and monitoring expertise.
About the job
Responsibilities
- Lead platform incident investigations, coordinating cross-functional teams through detection, mitigation, and resolution to minimize customer impact.
- Conduct post-incident root cause analyses across infrastructure, services, and cloud providers; identify systemic patterns and prevention measures.
- Design and implement customer-focused alerting pipelines and end-to-end observability workflows.
- Build automation tools, establish reusable monitoring patterns, and address reliability gaps affecting customer experience.
- Mentor junior engineers on observability patterns, alert design, and service health metrics.
- Participate in an on-call rotation.
Requirements
- At least 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or similar.
- Production experience with at least one major cloud provider: AWS, Azure, or Google Cloud.
- Proficiency with Docker and Kubernetes.
- Hands-on experience with monitoring, logging, and alerting tools such as ELK, Prometheus, Grafana, and PagerDuty.
- Ability to architect monitoring solutions that correlate metrics, logs, and traces.
- Strong proficiency in Python or a similar programming language, with the ability to build production-quality automation tools.
- Experience owning incident lifecycles from detection through resolution and post-mortem analysis in demanding production environments.
- Bachelor's, master's, or doctoral degree in Computer Science, Computer Engineering, or a related engineering field.
Skills
AWS, Azure, GCP, Docker, Kubernetes, Elk, Prometheus, Grafana, Pagerduty, Python, Observability, Incident Response
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.