PrismML builds high-performance foundation models designed to run efficiently across a wide range of environments—from edge devices to large-scale deployments. Our work spans models from ~1B to 100B+ parameters across LLMs, diffusion models, and other modalities, with a strong focus on scalable training, efficient inference, and real-world deployment. Today, that means a 27B-class model that runs on a phone in about 4 GB, down from roughly 54 GB at full precision.
Our Bonsai family of 1-bit and ternary models is designed to dramatically improve the efficiency of modern AI systems, enabling advanced intelligence to run with significantly lower memory usage, latency, and energy consumption across cloud and edge environments.
We are seeking a Staff-level (or higher) AI/ML engineer with expertise in multimodal systems to lead the development of capabilities that expand consumer use cases and product opportunities. This role focuses on building and integrating vision, speech, and other modalities into our core models, while providing technical leadership across AI/ML systems.
You will design, build, and integrate multimodal components optimized for performance, quality, and real-world deployment. Key responsibilities include:
You have a strong background in multimodal AI/ML systems and technical leadership:
You bring experience that aligns with building consumer-facing and efficient AI systems:
You enjoy turning advanced AI research into usable products, understand how multimodal systems unlock new user experiences, and think carefully about performance, efficiency, and deployment trade-offs. You take ownership of technical direction, work effectively across research and product teams, and actively support the growth of other AI/ML engineers.