Lead major reliability initiatives, mentor SRE engineers, and drive Infrastructure-as-Code and cloud modernization as a Staff Site Reliability Engineer on the SRE platform team. Requires 10+ years industry experience including software engineering, strong leadership, AWS, observability tools, and on-call participation.
148k – 193k/yr
Hybrid10+ YOEDevOps / SRE
About the role
What You’ll Do
Collaborate with platform architects and management to ensure reliability targets are met.
Advise engineering teams on best practices for measuring reliability and uptime.
Assist and respond to critical engineering incidents.
Lead and mentor SRE engineers to improve their engineering skills.
Provide technical guidance and best practices for use of cloud infrastructure and tooling. Be a driver for Infrastructure-as-Code within the platform.
Spearhead major reliability-focused initiatives and projects.
Help optimize our work to be customer-focused. Continually seek feedback from our customers on how we can improve.
Migrate legacy systems to modern, scalable cloud environments.
Help develop and drive a culture of continuous improvement with the Platform Engineering and Software Engineering groups.
Participate in an on-call rotation and occasionally act as Incident Commander.
The Skills and Experience You’ll Bring
Strong leadership and mentoring abilities, especially with SRE or Platform Engineering/Infrastructure teams.
Total of 10+ years industry experience.
3+ years of software engineering experience (JavaScript, Python, Go, Java/Kotlin, C++, etc).
Extensive experience with modern observability tools (Datadog preferred).
Extensive experience with cloud platforms (preferably AWS).
Demonstrated success in leading large technical initiatives, including design, project management and gaining executive buy-in.
Proven experience modernizing legacy code and infrastructure.
Ability to work closely with peer teams, platform/software architects and management to drive key reliability improvements.
Deep understanding of cloud infrastructure, automation, and best practices for reliability.
Experience with our tools (Kubernetes, ArgoCD, Terraform, Github Actions) a plus.
Why DAT?
Medical, Dental, Vision, Life, and AD&D insurance
Parental Leave
Up to 20 days of paid time off starting in year one
An additional 10 holidays of paid time off per calendar year
401k matching (immediately vested)
Employee Stock Purchase Plan
Short- and Long-term disability sick leave
Flexible Spending Accounts
Health Savings Accounts
Tuition Reimbursement Program
Employee Assistance Program
Additional programs - Employee Referral, Internal Recognition, and Wellness
Free TriMet transit pass (Beaverton Office)
Competitive salary and benefits package
Work on impactful projects in a cutting-edge environment
Collaborative and supportive team culture
Opportunity to make a real difference in the trucking industry
Employee Resource Groups
Skills
site reliability engineeringAWSKubernetesTerraformArgo CDGitHub ActionsDatadogPythonGoJavaInfrastructure As CodeObservability
Leads deployment, integration, and startup of BMS/EPMS/SCADA systems in data centers, ensuring seamless operation of HVAC, electrical, and monitoring infrastructure. Oversees contractors, troubleshoots protocols like BACnet/Modbus, and validates systems via FAT/SAT. Requires Bachelor's in engineering and hands-on data center automation experience.
148k – 170k/yr
On-siteDevOps / SRE
Staff Engineer, HMS Factory
Shield AIWashington, DC
Staff Engineer building automated testing frameworks, log analysis tools, and scenario-generation scripts on the Hivemind Platform SDK to enable rapid verification of AI Pilot behaviors. Requires strong Python expertise, simulation/testing experience, and ability to obtain SECRET clearance.
150k – 220k/yr
On-site7+ YOEDevOps / SRE
Staff Site Reliability Engineer
KongUnited States
Founding Staff SRE for Kong's internal developer platform (Volcano). Define reliability posture, build multi-region Kubernetes infrastructure, establish GitOps/CI-CD, and scale managed data services.
150k – 210k/yr
Remote7+ YOEDevOps / SRE
Staff Engineer, DevOps (4797)
Shield AISan Diego, CA +2
Owns and maintains C++ build systems for autonomous aircraft software, improves developer velocity by optimizing CI/CD pipelines, integrates testing with simulations, and implements monitoring to resolve issues quickly. Requires 7+ years experience with deep expertise in build tools and DevOps practices.
150k – 220k/yr
On-site7+ YOEDevOps / SRE
Staff Engineer, Software Integration (R4483)
Shield AISan Diego, CA +1
Integrates autonomy software stack for AI robotics platforms, including multi-agent systems, sensor processing, and hardware deployment across simulation, HIL, and flight environments. Requires 7+ years experience, Python/C++, CI/CD expertise, and strong systems integration skills.