Senior Platform Software Engineer
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
About the job
Responsibilities
- Ensure the reliability of core services, including PostgreSQL on RDS, ECS services, CI/CD build systems, and the OpenTelemetry observability stack on Datadog.
- Drive security initiatives that address risk with minimal friction.
- Create tools and share best practices that empower engineers to work effectively.
- Influence engineering standards for non-functional requirements.
- Collaborate across the company to understand requirements and share knowledge.
- Build the platform for AI development, including LLM tooling, engineering guardrails, and infrastructure for AI-powered product features.
- Lead incident response calmly and effectively, and model blameless postmortem and follow-up practices.
Requirements
- 6+ years of platform engineering or an adjacent background such as SRE or DevOps.
- Knowledge of AWS and open-source tools such as Terraform, Kafka, Temporal, PostgreSQL, and OpenTelemetry.
- Strong incident leadership, judgment, planning, communication, collaboration, and customer empathy.
Nice-to-haves
- Security and compliance experience in healthcare or another regulated industry, including HIPAA, SOC 2, or PCI.
- Experience with Datadog, Temporal, Kafka, Databricks, or Terraform.
- Experience with monorepo CI/CD systems at scale.
- Experience growing an organization at a Series B startup of approximately 30 engineers.
- Production experience with LLM tooling or infrastructure.
Compensation and Benefits
- Salary range: $180,000–$230,000 annually.
- Medical, dental, and vision coverage with nationwide coverage and virtual urgent care.
- Weekly therapy reimbursement up to $100.
- Up to 12 weeks of fully paid parental leave.
- 401(k), HSA, FSA, and monthly commuter benefits for NYC employees.
- Flexible paid time off.
- $100 monthly fitness stipend.
Skills
AWS, Terraform, Kafka, Temporal, Postgres, OpenTelemetry, Datadog, Databricks, ECS, Amazon Rds, CI/CD, Llm Tooling
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.
Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.