Staff Data Platform Engineer
As a Staff Data Platform Engineer, you will own the reliability, stability, and operational health of the data platform infrastructure. This role focuses on systems and infrastructure, ensuring proper deployment, monitoring, maintenance, and promotion across environments.
About the job
Responsibilities
- Own the reliability and availability of our data platform infrastructure across all environments
- Enforce and improve environment promotion discipline — staging is not prod, and prod is sacred
- Define and uphold SOPs around deployments, maintenance windows, and change management
- Instrument and monitor platform health using observability tooling; build alerting that means something
- Participate in architecture and deployment discussions; push back when something isn't ready
- Collaborate with data scientists, engineers, and product managers on infrastructure needs — as a partner, not an order-taker
- Identify and remediate reliability risks before they become incidents
- Support customer-facing and internal systems with a bias toward stability over velocity
Qualifications
The right candidate leans SRE. Data platform experience is additive — we will train the right person. Bullets marked with * are strongly preferred; all others are meaningful signal.
- Operational instinct — "the fear" — you've been burned by prod, you respect it, and you've built habits around it. You know what a proper maintenance window looks like, you communicate before you touch production, and you don't spin up new initiatives while something critical is still burning in.
- 3+ years in cloud infrastructure, SRE, or platform engineering (AWS preferred; GCP/Azure experience translates)
- High Availability architecture: blue/green deployments, data replication, load balancing
- Experience with workflow orchestration (Airflow or similar DAG-based schedulers — or general job scheduling/cron systems at scale)
- Strong Linux fundamentals and scripting (Bash, Python, or similar)
- Distributed data processing (Spark, PySpark, or similar big data frameworks — or experience managing clusters that run them)
- Containerization and orchestration (Kubernetes, Docker, or similar)
- Data ingestion, ETL, or streaming systems (Kafka, Flink, or similar — or experience operating message queues and pipelines)
- Infrastructure-as-code and provisioning (Terraform, Helm, or similar)
- OLAP and OLTP databases (Clickhouse, Postgres, Redshift, or similar — query patterns, indexing, and operational care)
- Monitoring, logging, and observability (Datadog, Prometheus, Kibana, or similar)
- Managed data platforms (Databricks or similar — administering and scaling, not just consuming)
- Network infrastructure fundamentals: load balancers, DNS, auto-scaling, multi-region topologies, proxies
- Security and access management: least-privilege, secrets management, controls for data systems
- MLOps concepts or tooling — a plus
What we value above technical skills
We are explicitly willing to trade depth in data tooling for the right operational character. Specifically:
- Humility — you don't know everything, you say so, and you ask before acting in unfamiliar territory
- Methodical execution — you minimize variables, you don't premature-optimize, you finish what you started before starting something new
- Communication — you tell the team what you're doing before you do it, especially in shared or production environments
- Ownership — when something goes wrong, you look inward first
- Independence – you can drive projects end-to-end, from ambiguous requirements to high quality deliverables. But you aren’t afraid to ask for help.
Benefits:
- Total compensation ($190,000 - $240,000)
- Equity compensation
- Health insurance coverage for you and your dependents
- 401K, FSA, and commuter benefits
- $150 monthly spending account
- $1,000 annual continued education benefit
- $500 Newbie Productivity Perk
- Unlimited PTO and sick days
- Monthly Company Wellness Day Off
- Snacks, drinks, and catered lunches at the office
- Team building events
Skills
AWS, Airflow, Bash, Python, Spark, Pyspark, Kubernetes, Docker, Kafka, Flink, Terraform, Helm, Postgres, Redshift, Datadog
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.