Senior Production Engineer
Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.
About the job
Responsibilities
- Own the health, resilience, and recovery of production systems.
- Design and build monitoring and observability platforms that reduce alert fatigue and accelerate root-cause analysis.
- Develop automation and self-healing capabilities for diagnosis and recovery workflows.
- Analyze incidents, identify systemic trends, and prevent recurring failures.
- Build reusable runbooks, diagnostic tooling, and recovery playbooks.
- Create reliable, standardized operational workflows for engineering teams.
- Partner with Platform Engineering on CI/CD pipelines, deployment safety, and infrastructure resilience.
- Champion Infrastructure as Code, GitOps, and SRE practices.
- Measure production health through SLIs, SLOs, and SLAs and use reliability data to guide priorities.
- Explore AI-assisted diagnostics and developer tooling.
Requirements
- Strong hands-on Python skills for automation and tooling.
- Experience in SRE, Production Engineering, Platform Engineering, or a related discipline with direct production ownership.
- Demonstrated experience building automation and diagnostic tooling that improves recovery times or reduces operational toil.
- Deep familiarity with cloud-native technologies, Kubernetes, containers, and distributed systems.
- Experience with observability platforms such as Datadog.
- Exposure to Infrastructure as Code with Terraform and GitOps deployment workflows using ArgoCD, GitHub Actions, or similar tools.
- Familiarity with Java, Go, Kafka, Redis, Snowflake, and PostgreSQL.
- Strong analytical, problem-solving, communication, and cross-functional collaboration skills.
- Product mindset for internal operational tooling, including usability, adoption, and documentation.
- Self-starter mentality and continuous-learning mindset.
Nice-to-have
- Fintech or financial industry experience.
Technology
- Kubernetes, AWS, Terraform, ArgoCD, GitHub Actions
- Kafka, Redis, PostgreSQL, Snowflake
- Datadog, Python, Go, Java
- gRPC, Protobuf, internal platform APIs, and developer tooling
Skills
Python, Kubernetes, AWS, Terraform, Argo CD, GitHub Actions, Datadog, Kafka, Redis, Snowflake, Postgres, Go, Java, gRPC, Protobuf
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Senior Platform Engineer responsible for architecting scalable, secure infrastructure and improving reliability, observability, and production operations. The role requires strong AWS, Infrastructure as Code, and Kubernetes experience, along with technical leadership and mentoring skills.
Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
Own and evolve a platform domain supporting reliable, secure, and cost-effective multi-tenant infrastructure. The role requires strong Kubernetes, cloud, Terraform, Helm, GitOps, Python, security, and agentic coding expertise, along with excellent technical judgment and communication.