Brimkerndocs

Documentation

Every part of the project, and how to use it: the chat, the embeddable SDK, the converter, the published models, the source. The side menu leads to each topic’s page.

Open the chat
The app itself: load a model and talk to it, entirely on your GPU.
Embeddable SDK
Put a local assistant on your own site with one script tag. Live demo included.
WGSL Terminal CLI
Run coding models on your GPU straight from your shell with Unix pipes.
GGUF → .brik converter
Repackage a model for streaming, in your browser. Nothing is uploaded.
Changelog
What changed, release by release, with the measurements behind each claim.
Published models
Our pre-quantized .brik models on Hugging Face, ready to stream.
Source code
The whole engine under MIT: WGSL kernels, the .brik format, the loaders, the app.

Getting started

Open the app and click the single button on the home screen. The model streams in once (149 MB for the default), is cached on your device, and every later visit starts in seconds: offline included. Nothing is ever uploaded: the weights come down to your machine and the computation happens on your GPU.

Requirements: a browser with WebGPU (Chrome, Edge, or Safari 18+). A discrete GPU helps for models above 1B parameters, but a laptop runs the small ones comfortably.

To go further: run any Hugging Face model, put the assistant on your own site, or use the terminal CLI.

Storage & offline

Model weights live in the browser cache, per site. The browser decides how much space it grants: often tens of gigabytes on a normal profile, but only ~1.5 GB in a private window, where a large model will not stay cached. The model browser warns you before a download that cannot fit.

Models you have not used for 30 days are cleaned up automatically (adjustable, or off, in the Storage panel). Conversations and locally converted .brik files are never touched.

Brimkern: open WebGPU engine, built by Romain Khanoyan. Local AI, WebGPU, on-device engines.