Documentation
Every part of the project, and how to use it: the chat, the embeddable SDK, the converter, the published models, the source. The side menu leads to each topic’s page.
Getting started
Open the app and click the single button on the home screen. The model streams in once (149 MB for the default), is cached on your device, and every later visit starts in seconds: offline included. Nothing is ever uploaded: the weights come down to your machine and the computation happens on your GPU.
Requirements: a browser with WebGPU (Chrome, Edge, or Safari 18+). A discrete GPU helps for models above 1B parameters, but a laptop runs the small ones comfortably.
To go further: run any Hugging Face model, put the assistant on your own site, or use the terminal CLI.
Storage & offline
Model weights live in the browser cache, per site. The browser decides how much space it grants: often tens of gigabytes on a normal profile, but only ~1.5 GB in a private window, where a large model will not stay cached. The model browser warns you before a download that cannot fit.
Models you have not used for 30 days are cleaned up automatically (adjustable, or off, in the Storage panel). Conversations and locally converted .brik files are never touched.
Brimkern: open WebGPU engine, built by Romain Khanoyan. Local AI, WebGPU, on-device engines.