Skip to content
Cerebras SystemsCerebras SystemsSunnyvale, CA

Hardware Analytics Engineer

Design and optimize hyperscale data pipelines and ML models for multi-terabyte hardware telemetry, performance analysis, anomaly detection, reliability forecasting, and efficiency optimization of AI servers and silicon hardware. Requires Master's degree and 3+ years experience.

214k – 225k/yr
Hybrid3+ YOEData Engineering

About the role

Job Duties

  • Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.
  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability.
  • Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.
  • Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.
  • Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.
  • Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.
  • Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.

Minimum Requirements

  • Master’s degree or foreign equivalent degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field and 3 years of experience as Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or a related occupation required.

Required Skills

  • Large-scale data pipeline architecture and ETL, distributed data processing (Hive, Spark), and dashboard development.
  • Python, SQL, Tableau, Linux, and automation scripting.
  • Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction.
  • Predictive modeling, statistical analysis, A/B testing, anomaly detection, and data visualization in hardware reliability and performance.
  • Hardware analytics for compute, storage, and AI servers; power and thermal optimization; GPU burn-in efficiency optimization; and reliability modeling for AI hardware systems and components including CPU, GPU, DRAM, and SSD.

Additional Information

Salary Range: $213,675.00 per year to $225,000.00 per year

Skills

PythonSQLTableauhiveSparkLinuxMachine LearningETLdata pipeline architecturepredictive modelingAnomaly DetectionA/B TestingStatistical Analysis
Hinge Health

Data Engineering Manager, Data & ML Platform

Hinge HealthSan Francisco, CA

As a Data Engineering Manager, you will lead the Data & ML Platform team, owning platforms for analytics, experimentation, and machine learning. You will guide the evolution towards a streaming-first, ML-ready architecture and partner with Data Science to operationalize models.

220k – 330k/yr
Hybrid5+ YOEData Engineering
Anyscale

Software Engineer (Ray Data)

AnyscaleSan Francisco, CA

Build, optimize, and scale Ray Data, a Python-native data processing engine for AI/ML workloads. Improve performance, ensure fault tolerance at high scale, and support production training and multi-modal batch inference for AI-native companies.

226k – 241k/yr
On-site4+ YOEData Engineering
GlossGenius

Software Engineer, Data Platform

GlossGeniusSan Francisco, CA

Designs and implements scalable data models, pipelines, and lakehouse infrastructure using Snowflake and Clickhouse to support analytics, ML, and products. Requires 5+ years data engineering experience, SQL/Python expertise, and leadership in data governance.

200k – 236k/yr
Hybrid5+ YOEData Engineering
OpenAI

Software Engineer, Distributed Data Systems (Sora)

OpenAISan Francisco, CA

Designs and scales distributed data infrastructure for large-scale multimodal training and evaluation at OpenAI. Collaborates with researchers to build reliable, high-performance systems handling massive data volumes in a fast-paced environment.

230k – 385k/yr
HybridData Engineering
OpenAI

Data Engineer, Analytics

OpenAISan Francisco, CA

Build and manage data pipelines and canonical datasets for product metrics, safety systems, and business decisions. Collaborate with cross-functional teams including Data Science and Research; requires 3+ years data engineering experience with Spark, ETL tools, and distributed systems.

230k – 385k/yr
On-site3+ YOEData Engineering