Site Reliability Engineer
Site Reliability Engineer improves, manages, and monitors production-critical infrastructure and data pipelines in a finance AI/ML firm. Collaborates on fault-tolerance, deployments, automation, and on-call incident response using Python, Linux, and cloud tools. Requires 2+ years experience and quantitative degree.
About the job
Responsibilities
- Improve fault-tolerance and maintainability of code in proprietary data pipelines and trading systems
- Diagnose and fix bugs in code
- Lead complex deployments
- Automate manual workflows
- Track and prioritize outstanding production-related issues
- Share an on-call rotation responding to incidents to ensure the continuous operation of production-critical systems
Requirements
- Experience with coding and debugging Python
- Experience with Linux
- Familiarity with Relational Databases & SQL
- Sharp analytical and problem-solving skills and a persistent drive to make things work (better)
- Strong growth mindset and a passion for learning
- Strong technical communication skills
- Attention to detail
- 2 years of relevant industry experience
- An undergraduate degree or comparable training in a quantitative field or equivalent, relevant industry experience
Preferred Qualifications
- Familiarity with best practices concerning code maintainability, documentation, quality assurance, continuous integration and deployment
- Experience supporting production systems
- Experience with any of the following: gRPC microservices, Postgres, Pandas, Golang, R, Git, Jenkins, Bazel, Prometheus, Grafana, Airflow, Kubernetes
Skills
Python, Linux, SQL, Kubernetes, Prometheus, Grafana, Postgres, pandas, Go, Airflow, Jenkins, Git, gRPC, Bazel, R
Similar jobs
DevOps / SRE jobsBuild Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.
Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.
Winter infrastructure and site reliability internship focused on building and operating minimal, on-premises backend infrastructure for a semiconductor fabrication facility. The role requires systems programming, Linux, networking, distributed systems, and hands-on infrastructure or automation experience.
Supports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.
Supports and evolves the networking, compute, Kubernetes, and ingress infrastructure powering PagerDuty’s real-time platform. Requires 0–1+ years of relevant experience, Linux production operations, cloud infrastructure knowledge, programming proficiency, and Infrastructure as Code experience.