Brimkerndocs

Models & the .brik format

What the engine loads and how: single-file GGUF straight from Hugging Face, supported architectures, quantization kernels, shareable test links, and the .brik streaming format with its in-browser converter.

Run any Hugging Face model

Brimkern reads single-file GGUF models directly: the format hosted by Hugging Face, with zero conversion or server compilation step. Paste any of these formats into the input field on the home screen, in the model browser, or pass them to the CLI:

Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF
https://huggingface.co/bartowski/DeepSeek-R1-Distill-Qwen-1.5B-GGUF
https://huggingface.co/unsloth/gemma-3-270m-it-GGUF/blob/main/gemma-3-270m-it-Q4_K_M.gguf
https://example.com/my-custom-model.gguf

The engine resolves the repository through the Hugging Face API, inspects available files, and picks the best quantization automatically (favoring .brik first, then Q4_K_M, Q4_K_S, Q4_0, Q4_1, Q5_K, Q3_K, Q8_0). Tokenizer, vocabulary, RoPE frequency, and context configurations are parsed directly from the GGUF header.

Supported architectures

Brimkern features dedicated, hand-crafted WGSL kernels matching modern model families:

Qwen (Alibaba)

Qwen 2, Qwen 2.5, Qwen 2.5 Coder, Qwen 3 (with <think> reasoning), and Qwen 3.5 SSM (DeltaNet causal conv + recurrent state).

Llama (Meta)

Llama 2, Llama 3, Llama 3.1, Llama 3.2, Granite 4.0 Micro, and Falcon 3. Interleaved RoPE pairs (ggml NORM) handled natively.

DeepSeek (DeepSeek AI)

DeepSeek-R1 Distill (Qwen & Llama backbones) from 1.5B to 7B with internal monologue decoding.

Gemma (Google)

Gemma 1, Gemma 2 (tanh logit softcaps 50/30, GELU gate, query pre-attn scaling), and Gemma 3 (alternating 5 local / 1 global sliding window attention).

Mistral (Mistral AI)

Mistral 7B and Ministral 3 with static YaRN frequency transform and attention scaling.

SmolLM (Hugging Face)

SmolLM, SmolLM2, and SmolLM3 featuring NoPE (1 in 4 layers without RoPE for long contexts).

Hybrides & Récurrents

Liquid AI LFM2/2.5 (gated short conv + attention) and BlinkDL RWKV-7 (constant ~1MB resident state replacing KV cache).

Quantizations & WebGPU kernels

Brimkern executes weight dequantization directly on the GPU in WebGPU compute shaders. Every kernel is verified against a CPU reference at startup (selfValidate) with graceful fallback and URL kill-switches.

Supported GGML / GGUF quantizations:

QuantizationBlock formatGPU kernelTypical use
Q4_K_M / Q4_K_S256 weights, 144 Bdequant_q4kDefault recommended choice for 1B-4B models
Q3_K (M, S, L)256 weights, 110 Bdequant_q3kFits 7B-8B models under ~3.8 GB VRAM on 8GB machines
Q4_032 weights, 18 Bdequant_q4_0Standard symmetric 4-bit quantization
Q4_132 weights, 20 Bdequant_q4_1Asymmetric 4-bit with min offset
Q5_0 / Q5_K32 / 256 weightsdequant_q5_0 / dequant_q5kHigher precision for critical projections
Q6_K256 weights, 210 Bdequant_q6kHigh-fidelity layers in mixed-precision GGUFs
Q8_032 weights, 34 Bdequant_q8_08-bit integer precision, near lossless output head
F16 / F321 weightdirect nativeUnquantized tensors (RMSNorm weights, biases)

Limits & exclusions

To preserve browser responsiveness and stability, some files and formats are intentionally rejected with clear diagnostic error messages:

  • Sharded GGUFs: Files named model-00001-of-00005.gguf are not supported. Brimkern streams contiguous HTTP byte ranges over a single file without loading multi-gigabyte blobs into system RAM. Use the single-file GGUF equivalent.
  • Standalone mmproj files: Files prefixed with mmproj-* contain image encoder vision projectors and cannot run standalone without a paired LLM.
  • Codebook IQ quantizations: Formats like IQ1_S, IQ2_XXS, IQ3_XXS, IQ4_XS rely on lookup tables that do not vectorize efficiently in direct WGSL shaders. Pick Q3_K, Q4_K_M, or Q4_0 instead.
  • Memory limits (VRAM): Models up to 1.5B (~1.2 GB) run smoothly on any mobile or laptop. Models from 3B to 4B (~2-2.6 GB) require 8GB+ system RAM. 7B models (~4-5 GB) require 16GB+ RAM.

The .brik format & converter

A .brik is a GGUF re-packaged for the browser: weights already quantized to int4/int8, laid out so each layer is one contiguous HTTP range, with the tokenizer embedded. The practical effect: the model loads by ranges (resumable, partially, genuinely offline afterwards) instead of as one multi-gigabyte download.

You can convert a GGUF yourself, in the browser. The file never leaves your machine: open the converter.

Brimkern: open WebGPU engine, built by Romain Khanoyan. Local AI, WebGPU, on-device engines.