Senior Software Engineer, Data Infrastructure
Senior Data Infrastructure Engineer responsible for building and operating reliable, low-latency streaming and batch data systems that support AI products. Requires 5+ years of production data infrastructure experience and expertise with technologies such as Kafka, Flink, ClickHouse, and Terraform.
About the job
Responsibilities
- Design and implement high-throughput data pipelines and streaming systems with strong SLOs, runbooks, and actionable telemetry.
- Build and operate real-time and batch ingestion infrastructure using Kafka, Flink, and Airflow.
- Own the analytical data layer, including schema design, query performance, and cost optimization across ClickHouse, BigQuery, or similar systems.
- Partner with research and product teams to architect data solutions, evaluate performance, and scale new features.
- Optimize pipeline and query latency through caching, partitioning, and data-path improvements to meet tight p95/p99 targets.
- Lead infrastructure-as-code and GitOps practices for data systems using Terraform, reusable modules, and policy-as-code.
- Participate in on-call and reduce operational toil through automation and elimination of recurring data issues.
Requirements
- 5+ years building and operating production data infrastructure at scale.
- Hands-on experience with ClickHouse, Kafka or equivalent messaging systems, and Flink or dbt.
- Proven ability to meet high-availability and low-latency targets across streaming and batch workloads.
- Strong observability and incident-response experience.
- Clear written communication and ability to turn ambiguous data requirements into reliable designs.
Nice-to-haves
- CDC experience with Debezium.
- Experience with Airflow, Dagster, or Prefect.
- Familiarity with Spark or Dask for large-scale data processing.
- Experience with Snowflake, BigQuery, Redshift, or Databricks.
- Early data, platform, or infrastructure engineering experience at another company.
- Strong Kubernetes experience with GKE, EKS, or AKS.
- Multi-cloud experience across Google Cloud, AWS, or Azure.
- Experience with customer-managed deployments.
Compensation and Benefits
- $200,000–$400,000 base salary plus equity.
- Medical, dental, and vision benefits.
- Life insurance and disability benefits.
- Retirement plan.
- Parental leave.
- Fertility and family-building benefits.
- Monthly wellness and lifestyle stipend.
- Daily office lunches and snacks.
- Flexible vacation policy.
Skills
Kafka, Apache Flink, Apache Airflow, ClickHouse, BigQuery, Terraform, GitOps, OpenTelemetry, Prometheus, Grafana, Datadog, Debezium, Kubernetes, Spark, dbt
Similar jobs
Data Engineering jobsBuild and operate petabyte-scale data infrastructure powering Discord’s insights and products. The role requires 5+ years of software engineering experience, strong programming skills, and experience with large-scale pipelines, streaming, orchestration, or data warehousing.
Build and lead the central data platform, covering ingestion, warehousing, orchestration, streaming, self-service frameworks, and trust layers. The role requires 5+ years of production data infrastructure experience, strong Python and SQL skills, and expertise with Snowflake and modern data tooling.
Build and operate large-scale revenue data pipelines powering billing and cost attribution, while improving reliability, latency, and correctness. The role requires strong Spark and Airflow experience, cross-functional problem-solving, and operational ownership of mission-critical production systems.
Build and operate low-latency systems that capture, normalize, and distribute real-time market data for institutional trading. The role requires backend engineering experience, Java or C++, market data infrastructure knowledge, and exchange connectivity expertise.
Build and own large-scale data models, batch and real-time pipelines, and data infrastructure that provide reliable datasets and insights across Plaid. The role requires 4+ years of data engineering experience, strong SQL and Python skills, and expertise with modern warehouses, lakes, and orchestration tools.