Full Stack LLM Engineer
This engineer brings up state-of-the-art language and AI models on Cerebras systems, working across model translation, compiler optimization, runtime integration, debugging, and performance tuning. The role requires broad AI toolchain experience, deep learning frameworks, C/C++, compiler development, and strong optimization skills.
About the job
Responsibilities
- Contribute to the end-to-end bring-up of machine learning models on Cerebras CSX systems.
- Work across the stack, including model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.
- Debug performance and correctness issues spanning model code, compiler intermediate representations, runtime behavior, and hardware utilization.
- Propose and prototype improvements across tools, APIs, and automation flows to accelerate future model bring-ups.
Requirements
- Bachelor's, master's, or PhD in Computer Science, Engineering, or a related field.
- Comfort navigating the full AI toolchain, including Python modeling code, compiler intermediate representations, and performance profiling.
- Strong debugging skills across performance, numerical accuracy, and runtime integration.
- Experience with deep learning frameworks such as PyTorch or TensorFlow, and familiarity with model internals including attention, mixture-of-experts, and diffusion.
- Proficiency in C/C++ programming and experience with low-level optimization.
- Proven experience in compiler development, particularly with LLVM and/or MLIR.
- Strong background in optimization techniques, particularly those involving NP-hard problems.
Compensation & Benefits
- Competitive salary and benefits package.
- Opportunities for professional growth and career advancement.
- Dynamic and innovative work environment.
- Opportunity to work on cutting-edge technologies and contribute to AI advancements.
Skills
Python, C, C++, PyTorch, TensorFlow, Llvm, Mlir, Compiler Development, Compiler Ir, Performance Profiling, Low-Level Optimization, Model Architecture, Graph Lowering, Runtime Integration, Numerical Accuracy
Similar jobs
Fullstack Engineering jobsProduct Engineer responsible for building Bolter’s AI agent runtime, workspace, generated applications, and production reliability systems. The role requires strong full-stack TypeScript experience, product judgment, rapid prototyping ability, and depth in AI systems, real-time infrastructure, or developer tooling.
Build and scale full-stack Creative and Studio products spanning APIs, user interfaces, infrastructure, and generative AI workflows. The role emphasizes product ownership, solving complex engineering problems, and expertise in Python, TypeScript, and React; no formal degree or experience requirement is specified.
Build and maintain user-facing features across a full-stack codebase, with opportunities to ship iOS functionality. The role requires proficiency in Go, Rust, or TypeScript, strong testing and communication practices, and ownership across the software development lifecycle.
Build customer-facing frontend experiences, data visualizations, and AI-powered interfaces for a conversation intelligence platform, while contributing to backend APIs. Requires 3+ years of production web application experience and strong frontend architecture skills.
Build Socket’s web application end-to-end, developing user interfaces, APIs, and data integrations while helping shape product direction and team foundations. The role requires production web application experience and proficiency with Node.js, JavaScript, React, TypeScript, and relational databases.