Skip to content
KustomerKustomerNew York, NY

Software Engineer, Cost Optimization

Own end-to-end cost optimization across AWS infrastructure and AI/LLM usage, building tooling, dashboards, and guardrails while balancing cost, performance, and reliability. Requires 4+ years of AWS infrastructure experience, cost-management expertise, and proficiency with Terraform and a programming language.

130k – 215k/yr
Hybrid4+ YOEDevOps / SRE

About the role

Responsibilities

Infrastructure & Cloud Cost Optimization

  • Own AWS cost optimization end to end, including EC2 right-sizing, Reserved Instances and Savings Plans, spot usage, storage tiering, and unused-resource cleanup.
  • Build dashboards and reporting to track infrastructure spend by team and service, and flag anomalies before they become expensive.
  • Partner with engineering teams to identify and eliminate waste, including idle resources, over-provisioned instances, and redundant environments.
  • Drive architectural decisions that balance cost, performance, and reliability.
  • Use Terraform to codify and enforce cost-conscious infrastructure standards.

AI Token Spend Optimization

  • Analyze and reduce AI/LLM token spend through model right-sizing, prompt efficiency, caching, and routing strategies.
  • Build developer tooling, guardrails, and observability around token usage.
  • Evaluate deterministic versus generative tool-selection tradeoffs from a cost perspective and recommend lower-cost approaches where appropriate.
  • Support AI capabilities by keeping them performant and cost-efficient.

Cross-Team Collaboration

  • Report cost trends and savings opportunities to engineering leadership.
  • Identify systemic inefficiencies and build tooling to prevent them from recurring.
  • Collaborate with FinOps and finance stakeholders to forecast and track spend against budget.
  • Drive alignment across engineering teams while maintaining development speed and momentum.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
  • 4+ years of experience with AWS infrastructure at scale, with a strong track record of reducing cloud costs.
  • Hands-on experience with AWS cost tools, including Cost Explorer, Cost and Usage Reports, and Trusted Advisor.
  • Experience with EC2 right-sizing and reservation strategies.
  • Proficiency with Terraform or other infrastructure-as-code tools.
  • Proficiency in a high-level programming language such as Go, Python, or JavaScript.
  • A cost-focused mindset, with an emphasis on measurable savings.

Nice to Have

  • Familiarity with LLM and AI infrastructure cost drivers, including token usage, model pricing, and caching strategies.

Technology Stack

  • AWS Cloud, including EC2 and AWS cost and billing tools
  • Terraform
  • MongoDB
  • Redis
  • Elasticsearch
  • ClickHouse
  • Kafka
  • Kinesis
  • Go
  • Python
  • JavaScript
  • React
  • Node.js

Compensation & Benefits

  • Competitive salary and stock options.
  • In the U.S.: 100% healthcare coverage, 401(k), WiFi and mobile reimbursement, and generous vacation policy.

Skills

AWSamazon ec2Terraformfinopsaws cost exploreraws cost and usage reportsaws trusted advisorKubernetesGoPythonJavaScriptMongoDBRedisKafkakinesis

Similar roles

DevOps / SRE jobs
Tulip

AI Enablement Engineer

TulipSomerville, MA

Build and maintain an internal agentic AI platform to accelerate developer workflows at Tulip. Identify high-impact AI opportunities in code generation, testing, debugging and tooling; own evals, standards, onboarding and measurement of AI adoption. Requires 5+ years software engineering experience with strong hands-on LLM/agentic AI and full-stack TypeScript skills.

130k – 180k/yrHybrid5+ YOEDevOps / SRE
Nominal

Software Engineer, Developer Infrastructure

NominalNew York, NY +3

Build, optimize, and maintain large-scale build systems (Bazel priority) and developer infrastructure including CI/CD, observability, and release automation for a fast-growing hardware-software platform company. Requires 4+ years experience with build systems at scale and large monorepos.

130k – 230k/yrOn-site4+ YOEDevOps / SRE
Onxmaps

Site Reliability Engineer III

OnxmapsBozeman, MT

Site Reliability Engineer responsible for deploying, monitoring, and maintaining highly available infrastructure on GCP using Terraform, Kubernetes, and various cloud services. Requires 5+ years experience (3+ in production), strong Kubernetes/IaC background, and on-call participation to ensure reliable systems for millions of users.

130k – 153k/yrHybrid5+ YOEDevOps / SRE
Kustomer

Software Engineer, Infrastructure

KustomerNew York, NY

Infrastructure Software Engineer building scalable backend systems, observability, and developer tools on the Foundation team. Lead projects on database sharding, event bus, search scaling, and latency; mentor engineers. Requires 5+ years with distributed systems, NoSQL (MongoDB), IaC (Terraform), and architecture ownership.

130k – 215k/yrHybrid5+ YOEDevOps / SRE
Trexquant

Linux Systems Engineer (USA)

TrexquantStamford, CT +1

Hands-on Linux Systems Engineer builds and maintains bare-metal servers, manages storage like ZFS, automates with Ansible and Bash, and ensures production reliability. Requires 3+ years Linux experience, physical server management, and on-call rotation with data center travel.

130k – 150k/yrOn-site3+ YOEDevOps / SRE