Sr. Software Engineer, DevOps
Builds and scales reliable infrastructure for SaaS applications using Kubernetes, Terraform, and GitHub CI/CD. Focuses on observability with Grafana/Prometheus, automation to reduce toil, production troubleshooting, and cross-team collaboration. Requires 5+ years Python experience.
About the job
What You'll Do
Infrastructure Management: Build, manage, and optimize infrastructure using Terraform, GitHub CI/CD, and Kubernetes.
Monitoring & Observability: Create visualizations and alerts that provide actionable insights using tools like Grafana, Prometheus/Mimir, OpenSearch, and Sentry.
Automation & Reliability: Identify manual or error-prone processes and replace them with automated, repeatable systems.
Production Troubleshooting: Diagnose and resolve production issues across application and infrastructure layers.
Documentation: Capture knowledge in runbooks, setup guides, and architecture diagrams to support operational maturity.
Collaboration: Partner with engineers across teams to drive adoption of DevOps and infrastructure best practices.
Scalability Planning: Help scale infrastructure and monitoring systems to meet growing demands.
Incident Participation: Participate in an on-call rotation and support incident response processes as needed.
Skills & Qualifications
Observability: Experience with metrics, logs, and traces using tools such as Grafana, Prometheus/Mimir, OpenSearch, Sentry, or similar.
Infrastructure as Code: Proficient with Terraform, Kubernetes, and containerization tools.
Programming Skills: 5+ years of experience with Python.
Linux Systems: Comfortable working with Linux-based environments and writing shell scripts.
Communication: Strong collaboration skills with a focus on asynchronous, written communication.
Documentation: Commitment to clear, comprehensive documentation and process standardization.
Initiative: Self-starter mindset with a proactive approach to solving operational challenges.
Version Control: Skilled in Git/GitHub-based workflows.
Nice to haves
- Cloud Experience: AWS (preferred), GCP, or Azure cloud infrastructure management.
- Networking Fundamentals: Familiarity with TCP/IP, DNS, routing, and load balancing concepts.
- Security: Understanding of cloud and infrastructure security best practices.
- Performance Tuning: Experience tuning application or infrastructure performance in production environments.
What We Offer
- Flexible paid time off (PTO)
- Expansive coverage for health, dental, and vision
- Employer contribution to Health Savings Accounts (HSA)
- Generous parental leave policy
- Full employee coverage for life insurance
- Home office stipend
- Cell phone/internet reimbursement
- Company-paid holidays
- 401(K) plan
Compensation
Based on market data and other factors, the salary range for this position is $180,000-$220,000 + Equity. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
Skills
Terraform, Kubernetes, Python, Grafana, Prometheus, Opensearch, Sentry, GitHub, CI/CD, Linux, AWS, Git
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.