Sr. Software Engineer, Platform Engineering
Platform engineer responsible for designing, building, and operating cloud infrastructure on AWS and Kubernetes. Focus on developer tooling, automation with AI, performance tuning, security, observability, and on-call support for a SaaS platform serving private capital markets. Requires 5+ years in DevOps/SRE/Platform roles.
About the job
Responsibilities
- Design, build and operate the cloud infrastructure that Chronograph runs on
- Maintain a first-class developer experience by building reliable internal tooling (e.g. contribute to PR environment platform allowing developers to spin up an entire application stack with its own production-scale, ZFS-based database clone)
- Build and use AI-assisted tooling to automate operational work, with sound judgment on safety and cost
- Monitor application performance and help developers debug production issues
- Improve the performance, scalability, and security of our infrastructure
- Participate in an on-call rotation, respond to production incidents with infrastructure assistance, and help teams pursue observability improvements during the post-mortem process
Requirements
- 5+ years of professional software engineering experience, primarily working in DevOps, SRE, Platform Engineering roles or similar
- Architected infrastructure and tools for production SaaS applications, particularly in a rapidly growing startup environment
- Experience working with public cloud providers like AWS, production Kubernetes clusters and managed database systems
- Developed and scaled thoughtful abstractions to help manage infrastructure as code, CI/CD workflows, and local development environments
- Tuned performance and hardened security posture across a variety of systems and databases, addressing issues related to networking and other configuration/performance issues
- Integrated observability tools that enable developers to own logging, metrics, tracing, alarming, etc. for their services
- The ability to evaluate different architectural approaches, make sound decisions given the pros and cons, and deliver a clear recommendation to the team
- Clear communication skills; able to help others understand complex technical issues and committed to writing high-quality documentation
- A passion for maintainable, readable, stable, secure, and scalable code, infrastructure, and processes; stay on top of emerging industry best practices and technologies
Nice-to-Haves
- Prior experience with Ruby on Rails and Node.js
- Our core services are implemented in Ruby on Rails and Node.js (prior experience is a nice-to-have)
Compensation and Benefits
- Salary Range (dependent on experience): $175,000—$215,000 USD
- Flexible work arrangements (including remote / hybrid)
- Competitive salary
- 401(k)
- Unlimited and flexible vacation
- Generous health benefits
- Team week events in HQ (Brooklyn, NY) three times annually for all employees
- Fully-paid parental leave
- ...and more!
Skills
AWS, Kubernetes, Ruby on Rails, Node.js, Infrastructure As Code, CI/CD, Observability, DevOps, SRE, Platform Engineering
Similar jobs
DevOps / SRE jobsLeads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.
Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.