5 days ago
TechCrunch Sep 17, 2026

PrismML hopes its tiny LLM will change how we all use AI

PrismML, a startup founded by Caltech researchers and led by professor Babak Hassibi, is pioneering a new approach to large language models (LLMs) by drastically reducing their size without sacrificing performance. The company recently launched Bonsai 2 27B, a compressed version of Alibaba’s Qwen3.8 27B model, shrinking it from around 60 GB to just 5.9 GB. This breakthrough allows such sophisticated AI models to run efficiently on personal computers and potentially even high-end smartphones, which could democratize access and shift the AI paradigm away from cloud dependency.

The compression technique behind PrismML’s models centers on reducing the bits required to store each model weight, using a “ternary” system with just three values: +1, -1, or 0. This method, unique to the company, enables a 9-10x decrease in memory use while maintaining nearly 98% of the original model's benchmark performance. This is a significant improvement from their initial version, Bonsai 1, which achieved 95% parity and has been downloaded over 11 million times, indicating strong market interest in smaller yet capable LLMs.

PrismML’s ambitions extend to applying their compression technology to much larger models with several hundred billion parameters. Hassibi expects that larger models will be even easier to compress while retaining their intelligence because the abundance of parameters offers more room for efficient reduction. The company’s advisor Ion Stoica, a notable figure linked to Databricks and Berkeley’s Sky Computing Lab, highlights the privacy and cost benefits of running advanced AI directly on users’ devices without cloud reliance, potentially transforming how AI services are accessed.

The startup has so far raised $22.25 million in a seed round, supported by investors including Khosla Ventures, Cerberus Capital, and Caltech. While PrismML is not alone in the field of LLM compression—companies like Multiverse Computing also compete—they differentiate themselves by delivering highly compressed models with minimal loss in accuracy, signaling a promising future for AI that’s accessible, private, and device-friendly.

0
0 Read source
Share this post
Facebook Twitter LinkedIn

Discussion

0 comments

No comments yet

Start the discussion with a take, question, or market read.