PrismML builds high-performance foundation models designed to run efficiently across a wide range of environments—from edge devices to large-scale deployments. Our work spans models from ~1B to 100B+ parameters across LLMs, diffusion models, and other modalities, with a strong focus on scalable training, efficient inference, and real-world deployment. Today, that means a 27B-class model that runs on a phone in about 4 GB, down from roughly 54 GB at full precision.
Our Bonsai family of 1-bit and ternary models is designed to dramatically improve the efficiency of modern AI systems, enabling advanced intelligence to run with significantly lower memory usage, latency, and energy consumption across cloud and edge environments.
We are seeking a Senior-level (or higher) AI/ML engineer with deep expertise in systems and kernel development to lead efforts in optimizing low-level performance across our model stack. This role focuses on designing and implementing high-performance kernels that accelerate inference and training for highly efficient model architectures, including 1-bit and other compressed representations, across diverse hardware platforms.
You will design, implement, and optimize high-performance kernels and low-level systems to maximize efficiency across a range of inference runtimes and hardware targets. Core responsibilities include:
You bring deep experience in systems engineering and performance optimization for AI/ML workloads:
You have additional experience aligned with high-performance AI systems and hardware-aware optimization:
You have built and optimized kernels or low-level systems that significantly improve performance for large-scale AI workloads. You understand how model architecture, numerical representation, and hardware interact, and you know how to push systems to their limits. You care deeply about efficiency and performance, enjoy working close to the hardware/software boundary, and are comfortable leading efforts that span models, runtimes, and infrastructure.