Skip to content
NinjaTraderNinjaTraderChicago, IL

Operational Resiliency Engineer

Provide front-line operational support and incident response for critical FCM financial processes at NinjaTrader. Design monitors, maintain runbooks and DAGs, implement observability with OpenTelemetry/Kafka, and collaborate with engineering to improve resiliency of distributed systems.

100k – 150k/yr
Hybrid2+ YOEDevOps / SRE

About the role

What you’ll do

NinjaTrader is looking for an L1 Operational Resilience Specialist to provide front-line operational support to our FCM Core team across Account Acquisition, Compliance/AML, and Treasury functions. This role sits at the intersection of engineering and operations — you’ll keep critical financial processes running smoothly while working closely with Engineering, Product, QA, and cross-functional partners to drive continuous improvements in reliability and observability.

In this role you will:

  • Manage incident response by acknowledging and triaging incoming events, applying runbooks to restore service quickly and consistently
  • Design and implement metrics and monitors across our distributed application, collaborating with the SRE team to integrate into our observability framework
  • Develop and continuously improve runbooks for FCM processes, ensuring clear, repeatable procedures for operational events
  • Understand and maintain the DAG controlling all FCM processes and their dependencies, ensuring accurate sequencing and reliable execution
  • Implement robust monitoring and logging using industry-standard observability tooling such as OpenTelemetry
  • Update and maintain technical documentation describing the critical path processing for FCM; establish SLA, SLO, and SLI targets for each process
  • Manage event platform systems such as Kafka and IBM MQ to ensure reliable data connectivity with external partners
  • Surface operational feedback to engineering teams to drive improvements in the resiliency and robustness of business processes
  • Leverage AI tooling to proactively identify gaps in monitoring, analysis, documentation, and other operational domains

What you’ll need

  • BS or MS in Computer Science, Software Engineering, or a related field — or equivalent practical experience
  • 2+ years of experience working with distributed systems in a production environment
  • Comfort leveraging AI tooling to improve workflows and surface operational insights
  • Strong understanding of security best practices for backend services
  • Excellent verbal communication skills with a collaborative, team-first approach

Bonus Points for

  • Familiarity with finance back office systems and trading concepts
  • Background in multi-tenant SaaS platform architectures
  • Experience working in a regulated fintech, trading, or brokerage environment

Compensation

The salary range for this role will be $100,000.00 - $150,000.00 USD. In addition, this position will also receive an annual target bonus of 10%. Bonus pay at NinjaTrader is based on individual performance (50%) as well as company/team performance (50%).

Salary and bonus earnings are only two components of the total compensation package offered by NinjaTrader. NinjaTrader offers a 401(k) plan through ADP under which the company will match up to 3.5% of employee contributions. Annual paid time off allowance accrues at a rate of 18 days per year plus seven paid holidays.

Skills

Distributed SystemsIncident ResponseObservabilityOpenTelemetryKafkaibm mqrunbooksslaslosliAI Toolssecurity best practicesbackend services

Similar roles

DevOps / SRE jobs
9 Mothers

IT Engineer

9 MothersAustin, TX

DevOps Engineer building automated CI/CD pipelines, AI-powered internal tooling and agents, and fleet deployment systems for a defense startup developing counter-drone AI hardware. Requires 2+ years DevOps experience, strong Linux and networking skills, and US citizenship for ITAR compliance.

100k – 140k/yrOn-site2+ YOEDevOps / SRE
PointOne

Product Reliability Engineer

PointOneNew York, NY

Owns end-to-end system reliability, incident response, observability, and proactive stability improvements in a serverless AWS environment. Requires 2+ years software engineering with production-facing experience, strong debugging, and hands-on AWS/Go/TypeScript skills.

100k – 160k/yrOn-site2+ YOEDevOps / SRE
PointClickCare

Intermediate AI-Enabled DevOps Engineer

PointClickCareUnited States

Builds and operates cloud infrastructure and CI/CD pipelines for AI-enabled workloads, focusing on automation, reliability, and containerized deployments in Kubernetes. Requires 2-4+ years DevOps experience, IaC, scripting, and cloud providers like AWS/Azure/GCP.

106k – 118k/yrRemote2+ YOEDevOps / SRE
MongoDB

Cloud Operations Engineer

MongoDBUnited States

Cloud Operations Engineer on 2nd shift weekends responsible for monitoring Atlas platform, diagnosing incidents, on-call rotations, automation, and ensuring uptime for MongoDB customers in FedRamp environments. Requires 2+ years DevOps/SRE experience, Linux expertise, cloud familiarity, and scripting skills.

90k – 176k/yrOn-site2+ YOEDevOps / SRE
Fab2

Infrastructure Software Engineering Intern - Fall

Fab2San Francisco, CA +1

Infrastructure Software Engineering Intern designs, builds, and manages on-premises backend infrastructure for a semiconductor fab, focusing on bare-metal Linux, automation, monitoring, and reliability using Rust and Go. Requires strong systems programming skills and a portfolio of real projects; pursuing BS in CS/CE.

110k – 132k/yrOn-siteEntry levelDevOps / SRE