Site Reliability Engineer
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
About the job
Responsibilities
- Serve as the first responder for P1/P2 production incidents, triaging and stabilizing infrastructure-level issues within defined response SLAs.
- Diagnose failures primarily from system logs across Kubernetes, RabbitMQ, and PostgreSQL, identifying the affected component before escalation.
- Distinguish infrastructure failures from application or business-logic failures.
- Identify and submit fixes for infrastructure-level issues; escalate application or business-logic issues to the platform team with clear context.
- Participate in an on-call rotation, including off-hours and alternating weekend coverage.
- Communicate incident status clearly to stakeholders and provide clean handoffs to the resolution-owning team.
- At more senior levels, harden systems, drive scaling and capacity work, and propose structural improvements.
Requirements
- Hands-on enterprise experience with Kubernetes, RabbitMQ, and PostgreSQL, ideally in financial services.
- Strong working knowledge of Microsoft Azure.
- Experience troubleshooting production systems under time pressure.
- Sound judgment when assessing severity, escalation, and resolution paths.
- Ability to diagnose unfamiliar systems primarily from logs rather than source code.
- Clear, calm communication during live incidents.
Nice-to-haves
- Experience with Google Cloud or AWS.
Compensation and Schedule
- Hourly, contract-based engagement; not eligible for bonus or equity compensation.
- Minimum 10 hours per week during regular business hours.
- Rotational weekend coverage of 8 hours on Saturday and 8 hours on Sunday every other weekend.
- Approximately 72 or more hours per month total.
- Compensation is adjusted by location, market conditions, cost of living, experience, skills, and internal pay equity.
Skills
Kubernetes, RabbitMQ, Postgres, Microsoft Azure, GCP, AWS, Incident Response, Production Troubleshooting, Log Analysis
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Build and operate internal developer platforms that improve engineering velocity, reliability, and security. The role spans developer tooling, CI/CD, GitOps, AI-assisted development, and automated engineering guardrails.