Senior Platform Engineer
Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.
About the job
Responsibilities
Platform and Cloud Infrastructure
- Own and evolve AWS infrastructure across compute, networking, storage, and managed services.
- Design and maintain highly available infrastructure with predictable performance and financial correctness.
- Lead platform architecture decisions, service migrations, and runtime changes, such as Redis to Valkey and EKS to ECS/Fargate.
- Balance reliability, cost, and operational simplicity when making infrastructure decisions.
Deployment, Reliability, and Operations
- Design and maintain safe, repeatable, and observable deployment pipelines.
- Own reliability through capacity planning, failure modeling, and controlled change management.
- Lead incident response and root-cause analysis for infrastructure failures.
- Participate in on-call rotations and improve operational processes.
Observability, Security, and Governance
- Build and maintain observability across infrastructure and services, including metrics, logs, tracing, and alerting.
- Secure AWS resources through appropriate IAM policies, secrets management, and network boundaries.
- Identify and address infrastructure risks involving scale, cost, and security.
Technical Leadership and Collaboration
- Partner with application engineers to communicate platform capabilities and constraints.
- Implement infrastructure changes hands-on.
- Establish infrastructure, deployment, and operations standards as the team grows.
- Mentor platform engineers and improve organizational operational maturity.
Requirements
- 8+ years of experience building and operating production infrastructure in cloud environments.
- Deep experience with AWS core services, including EC2, ECS/EKS, VPC, IAM, RDS, ElastiCache, ALB/NLB, and CloudWatch.
- Strong understanding of containerized workloads and orchestration tradeoffs.
- Experience designing systems for high availability, fault tolerance, and controlled failure.
- Hands-on infrastructure-as-code experience with Terraform, CloudFormation, or an equivalent tool.
- Experience planning and executing infrastructure migrations safely.
- Experience debugging production incidents involving networking, scaling, or service degradation.
Nice-to-haves
- Experience operating infrastructure for financial systems or other high-reliability domains.
- Experience scaling infrastructure during periods of rapid growth.
- Strong observability and operational hygiene practices informed by past incidents.
- Experience simplifying, retiring, or rearchitecting over-engineered infrastructure.
- Comfort working with a Rails-based application stack while remaining tool-agnostic.
Compensation and Benefits
- Salary: $190,000–$200,000, depending on experience, plus benefits.
- Generous paid time off and company holidays.
- Company-paid short-term disability.
- Employer-covered health and dental insurance for direct employees; higher-tier and dependent coverage may involve additional employee cost.
- Vision plan available at additional employee cost.
- Child care benefits and parental leave.
Skills
AWS, Amazon Ec2, Amazon Ecs, Amazon Eks, Amazon Vpc, Aws Iam, Amazon Rds, Amazon Elasticache, Terraform, CloudFormation, Kubernetes, CloudWatch, Infrastructure As Code, Observability, Ruby on Rails
Similar jobs
DevOps / SRE jobsBuild and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.
Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.