Junior Software Engineer
Junior Site Reliability Engineer supporting production operations, observability, incident response, and automation for critical services. The role suits candidates with 0–2 years of experience, programming or scripting skills, and an interest in cloud infrastructure and distributed systems.
About the job
Responsibilities
Production Engineering
- Support operation of critical services owned by the SRE team.
- Assist with monitoring service health, availability, and performance.
- Participate in incident response activities under mentorship.
- Learn on-call practices and production operations.
- Assist with deployment validation and operational readiness activities.
Reliability Engineering and Observability
- Support implementation of monitoring, alerting, dashboards, and observability improvements.
- Use logs, metrics, traces, and events to help diagnose production environments.
- Identify monitoring gaps and opportunities to improve visibility.
- Participate in post-incident reviews and reliability improvement initiatives.
- Develop foundational understanding of SLIs, SLOs, and error budgets.
Automation and AI Operations
- Develop scripts and automation to eliminate repetitive operational work.
- Contribute to internally developed AI-powered operational tools.
- Support AI-assisted runbooks and operational workflows.
- Identify opportunities to reduce toil through software engineering and automation.
Engineering Enablement
- Partner with development teams to understand production systems.
- Assist teams in adopting reliability and observability best practices.
- Support operational readiness reviews.
- Learn how reliability influences software design and architecture decisions.
Requirements
- 0–2 years of experience through employment, internship, or co-op programs.
- Degree or diploma in Computer Science, Software Engineering, or equivalent experience.
- Basic programming or scripting experience.
- Strong problem-solving skills.
- Curiosity and willingness to learn complex systems.
- Effective communication and teamwork skills.
Nice-to-Haves
- Exposure to Azure cloud services.
- Exposure to Kubernetes or containerized applications.
- Familiarity with observability platforms such as Datadog, AppDynamics, Grafana, Prometheus, or ELK.
- Interest in AI, automation, and distributed systems.
Skills
Site Reliability Engineering, Programming, Scripting, Azure, Kubernetes, Containers, Datadog, Appdynamics, Grafana, Prometheus, Elk, Observability, Distributed Systems, Ai Operations
Similar jobs
DevOps / SRE jobsBuild and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.
Build and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.