Models & the .brik format
What the engine loads and how: single-file GGUF straight from Hugging Face, supported architectures, quantization kernels, shareable test links, and the .brik streaming format with its in-browser converter.
Run any Hugging Face model
Brimkern reads single-file GGUF models directly: the format hosted by Hugging Face, with zero conversion or server compilation step. Paste any of these formats into the input field on the home screen, in the model browser, or pass them to the CLI:
Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF https://huggingface.co/bartowski/DeepSeek-R1-Distill-Qwen-1.5B-GGUF https://huggingface.co/unsloth/gemma-3-270m-it-GGUF/blob/main/gemma-3-270m-it-Q4_K_M.gguf https://example.com/my-custom-model.gguf
The engine resolves the repository through the Hugging Face API, inspects available files, and picks the best quantization automatically (favoring .brik first, then Q4_K_M, Q4_K_S, Q4_0, Q4_1, Q5_K, Q3_K, Q8_0). Tokenizer, vocabulary, RoPE frequency, and context configurations are parsed directly from the GGUF header.
Supported architectures
Brimkern features dedicated, hand-crafted WGSL kernels matching modern model families:
Qwen 2, Qwen 2.5, Qwen 2.5 Coder, Qwen 3 (with <think> reasoning), and Qwen 3.5 SSM (DeltaNet causal conv + recurrent state).
Llama 2, Llama 3, Llama 3.1, Llama 3.2, Granite 4.0 Micro, and Falcon 3. Interleaved RoPE pairs (ggml NORM) handled natively.
DeepSeek-R1 Distill (Qwen & Llama backbones) from 1.5B to 7B with internal monologue decoding.
Gemma 1, Gemma 2 (tanh logit softcaps 50/30, GELU gate, query pre-attn scaling), and Gemma 3 (alternating 5 local / 1 global sliding window attention).
Mistral 7B and Ministral 3 with static YaRN frequency transform and attention scaling.
SmolLM, SmolLM2, and SmolLM3 featuring NoPE (1 in 4 layers without RoPE for long contexts).
Liquid AI LFM2/2.5 (gated short conv + attention) and BlinkDL RWKV-7 (constant ~1MB resident state replacing KV cache).
Quantizations & WebGPU kernels
Brimkern executes weight dequantization directly on the GPU in WebGPU compute shaders. Every kernel is verified against a CPU reference at startup (selfValidate) with graceful fallback and URL kill-switches.
Supported GGML / GGUF quantizations:
| Quantization | Block format | GPU kernel | Typical use |
|---|---|---|---|
| Q4_K_M / Q4_K_S | 256 weights, 144 B | dequant_q4k | Default recommended choice for 1B-4B models |
| Q3_K (M, S, L) | 256 weights, 110 B | dequant_q3k | Fits 7B-8B models under ~3.8 GB VRAM on 8GB machines |
| Q4_0 | 32 weights, 18 B | dequant_q4_0 | Standard symmetric 4-bit quantization |
| Q4_1 | 32 weights, 20 B | dequant_q4_1 | Asymmetric 4-bit with min offset |
| Q5_0 / Q5_K | 32 / 256 weights | dequant_q5_0 / dequant_q5k | Higher precision for critical projections |
| Q6_K | 256 weights, 210 B | dequant_q6k | High-fidelity layers in mixed-precision GGUFs |
| Q8_0 | 32 weights, 34 B | dequant_q8_0 | 8-bit integer precision, near lossless output head |
| F16 / F32 | 1 weight | direct native | Unquantized tensors (RMSNorm weights, biases) |
Limits & exclusions
To preserve browser responsiveness and stability, some files and formats are intentionally rejected with clear diagnostic error messages:
- Sharded GGUFs: Files named model-00001-of-00005.gguf are not supported. Brimkern streams contiguous HTTP byte ranges over a single file without loading multi-gigabyte blobs into system RAM. Use the single-file GGUF equivalent.
- Standalone mmproj files: Files prefixed with mmproj-* contain image encoder vision projectors and cannot run standalone without a paired LLM.
- Codebook IQ quantizations: Formats like IQ1_S, IQ2_XXS, IQ3_XXS, IQ4_XS rely on lookup tables that do not vectorize efficiently in direct WGSL shaders. Pick Q3_K, Q4_K_M, or Q4_0 instead.
- Memory limits (VRAM): Models up to 1.5B (~1.2 GB) run smoothly on any mobile or laptop. Models from 3B to 4B (~2-2.6 GB) require 8GB+ system RAM. 7B models (~4-5 GB) require 16GB+ RAM.
Instant test links
Any model can be turned into a link that loads it directly: handy to share a demo, to file a bug report, or to point a colleague at an exact quantization.
https://brimkern.com/chat?model=Qwen/Qwen3-0.6B-GGUF https://brimkern.com/chat?model=Qwen/Qwen3-0.6B-GGUF&file=Qwen3-0.6B-Q8_0.gguf https://brimkern.com/chat?gguf=https://example.com/model.gguf https://brimkern.com/chat?brik=https://example.com/model.brik
?model= resolves the repository through the Hub API and picks the best loadable file (a .brik wins over a GGUF). ?file= forces one exact quantization. ?gguf= and ?brik= take a direct URL, for models you host yourself.
The .brik format & converter
A .brik is a GGUF re-packaged for the browser: weights already quantized to int4/int8, laid out so each layer is one contiguous HTTP range, with the tokenizer embedded. The practical effect: the model loads by ranges (resumable, partially, genuinely offline afterwards) instead of as one multi-gigabyte download.
You can convert a GGUF yourself, in the browser. The file never leaves your machine: open the converter.
Brimkern: open WebGPU engine, built by Romain Khanoyan. Local AI, WebGPU, on-device engines.