Brimkern · Changelog
WebGPU-accelerated LLM inference, 100% in your browser.
Brimkern CLI & Native Dawn Engine: on-device terminal inference with zero browser overhead via Google Dawn WebGPU bindings (25× faster boot measured on Apple Metal: 0.93s vs 23.46s), Claude Code / Gemini-style workflows (@file injection, git shortcuts), and dedicated coding models (Qwen 2.5 Coder 0.5B & 1.5B).
In-process Native Dawn WebGPU Engine
- Replaced headless Chromium requirement with direct Google Dawn N-API bindings. Startup latency drops from 23.46s under Playwright Chromium to 0.93s on Apple Metal (25× speedup), and idle memory footprint falls from ~350 MB to ~45 MB RAM. 100% of WGSL compute shaders run unchanged on physical hardware.
- Persistent disk caching for HTTP Range tensor requests in ~/.cache/brimkern/ranges/. The first download warms local disk blocks; subsequent runs achieve 18/18 cache hits in 0ms with zero network requests.
- Automatic dual-runtime fallback: native Dawn executes in-process by default, with seamless fallback to headless Chromium for arbitrary full-GGUF dequantization or environments without native binaries (--native and --chromium flags).
New Developer Models & Context Workflows
- Added Qwen 2.5 Coder 0.5B Instruct (491 MB GGUF Q4_K_M, marked mobile: true in the catalog presets) and Qwen 2.5 Coder 1.5B Instruct (1.12 GB) alongside our lightweight 149 MB LFM2.5 Coder and RWKV-7 G1a 0.4B RNN.
- Claude Code and Gemini CLI workflow tools: file context injection (@path/to/file:start-end with sensitive secret guardrails), git shortcuts (/diff, /commit conventional proposals, /review <file>), /copy system clipboard bridge, and /stats token speed & $0.00 cost tracker.
SDK 0.3.0: the widget runs RWKV-7 .brik models — the guard that said “LFM2 only” was the last thing missing, the whole RWKV engine was already in the bundle. Built to answer a product question by measurement; the answer closed one path and opened a precise trade-off.
RWKV-7 in the widget
- model now accepts an RWKV-7 .brik URL alongside LFM2. The dispatch mirrors the app’s: the recurrent engine (WGSL kernels, drivers, World tokenizer path) was already shipped inside sdk.js — only the pure model class and one guard stood between integrators and Apache-licensed models. Prompts switch to RWKV’s plain-text User:/Assistant: format, and generation now also stops on a textual turn marker: the World vocabulary has no special turn tokens, so without that cut a small model happily writes both sides of the conversation. The bundle grows by 15 KB.
- Nothing moved for the default model: RAG 6/6, dialogue 33/33 over three rounds (only 2 caught by the safety net — the model earned 31 alone), API surface 50/50.
Two measurements that settle a product question
- The lightweight-widget question (“can the default model weigh ~100 MB instead of 149?”) had been blocked on a model decision since mid-August. The dispatch made it measurable. RWKV-7 G1 0.1B (128 MB, Apache): 6/24 on the widget’s RAG cases — it serves its canonical refusal while the right document sits selected under its eyes, and when it does read, it copies the whole document, forbidden number included. Reading documents is simply above what a 0.1B can do; that path is closed whatever the file format.
- RWKV-7 G1a 0.4B (304 MB, Apache): 10/12 — refusals, greetings and same-document number disambiguation all hold; the two failures are the same single case, reading a row out of a size table (it answers 26.0/26.5 cm instead of 27.0). The licence trade-off is now a figure, not a feeling: the Apache option costs twice the download and loses table reading against the 149 MB default (12/12, LFM 1.0 licence).
SDK 0.2.1: classification and one-shot generation join the resident GPU path — ×20 measured on the /local-ai demo, and video prompt enrichment rides along. Plus a lock on the shared GPU state: two widgets on one page can no longer silently corrupt each other’s answer.
The pure class joins the resident path
- The widget’s chat has run on the resident path since 0.1.x — chunked prefill on the GPU, one ~512-byte readback per token. But classify() and the direct callers of generate() (the /local-ai demo, the video prompt enricher) still paid the historical JS forward: about ten GPU submissions and a full-vocabulary readback PER TOKEN OF PROMPT. Classification is 100% prompt reading, which made it the single worst case of the whole engine.
- Measured with a new benchmark (alternating arms on the same page, the JS path as its own control via a kill-switch, and the OUTPUT checked on every shot — a speedup that changes the answer is not a speedup): classification 4.2 s → under 0.2 s (×20), extraction 2.8 s → under 0.2 s (×14). “Under 0.2 s” is the measurement floor of the harness itself, so both ratios are lower bounds. The answers are identical in both arms.
- Nothing changes for integrators: same API, same answers, and the JS path remains as an automatic fallback wherever the resident path is unavailable. The existing benchmarks: RAG 24/24 (two rounds, both languages), dialogue 33/33, API surface 50/50. Tools measured 14/15 with a failure that moves from round to round — and the frozen 0.2.0 bundle, untouched by this change, fails the same case the same way: sampling variance of a 230M that ignores its injected fact one draw out of several, not a regression of the port.
A lock on the shared GPU state
- The recurrent GPU state (attention K/V, conv state) is a single slot per engine — and the engine is a singleton per model URL, shared by every widget and session on the page. Two generations running at once would silently steal that slot from each other and both produce invalid output. Resident entry points are now serialized per model: the second caller waits its turn instead of corrupting the first. Nothing prevented this before; it simply had not happened yet on a page with one widget.
SDK 0.2.0: the widget gets tools — arithmetic, today’s date, your own functions — and stops imposing its look: theme (light, dark, or following the visitor’s system), corner, size, and every label can now be yours. The README had been promising “Tools are next” since 0.1.4; here they are, designed the only way that measurably works at this model size.
Tools: the model never decides, and that is the point
- tools: ['calc', 'date', { name, match, run }] on embed() and createSession() alike. The design is the one the app already validated in production: below ~3B parameters, letting the model emit tool calls produces hallucinated calls — so no decision is ever delegated to it. Detection is deterministic (a regex or your predicate), your function runs in your page, and the model receives the result as a bracketed fact in the turn, exactly like a knowledge note. A 230M is notoriously bad at arithmetic; with 'calc' it answers 127×9 with the exact figure, because the figure is handed to it.
- Three findings from the benchmark, each fixed by a lesson this engine already knew. A date line in the system prompt is not enough: asked “What year is it?”, the model served its trained refusal (“my knowledge is based on data up until July 2024”) without reading the line two sentences above — a bracketed note on date questions fixed it (“It is 2026.”). A custom tool got “I don’t have access to real-time inventory data” right next to a note containing precisely the stock asked for: the guardrail sentence literally said “You have no tools”, which became false the moment tools were declared — it is now conditional. And the form has to be SHOWN: two pinned examples demonstrate copying the bracketed fact, fabricated by the same function that builds the real turns, so example and prompt cannot drift apart.
- The contract protects the visitor: a tool that throws, hangs (10-second cap) or returns nothing is simply absent from the turn — the turn itself never fails because of a tool, the same rule as event listeners. Results are capped at 600 characters: a result is a fact, not a report. A tool event announces each result before generation, the traceability twin of onSources. And everything is opt-in: without a tools field, not one byte of the measured prompts changes.
- Measured on the real model, three rounds: 15/15 — the exact figure for arithmetic, the current year, a page-provided stock count reaching the answer, and the two interactions with knowledge documents that had to hold: the off-notes refusal stays intact when no tool triggers, and a calculation wins over the refusal instruction when one does. The existing benchmarks did not move: RAG 12/12 in both languages, dialogue 33/33, API surface 50/50, plus 24 unit assertions on detection and guardrails that run without a browser.
The widget stops imposing its look
- theme: 'light' | 'dark' | 'auto' — auto follows the visitor’s prefers-color-scheme, live: switch the OS and the widget follows. The default stays light, the exact look every current integrator validated on their page. position anchors it bottom-right or bottom-left; width and height are clamped numbers (300–480 × 380–720), and small screens still cap to the viewport. Everything lands as per-element custom properties — the rule learned with the accent color: a shared stylesheet must contain nothing widget-specific, or the second widget silently inherits from the first. Two widgets on one page can now differ in theme, corner and colour, and it is checked in a real browser.
- labels overrides any of the widget’s strings — launcher and close button, placeholder, the “Local AI” note, error prefix, the fallback and help phrases, the sources label, the loading phases. lang covers English and French out of the box; labels is the door to every other language, or simply to your own tone. Labels are rendered as text, never as markup, and theme takes a keyword, not CSS: nothing an integrator passes can inject styles or markup into the host page — the safeAccent rule, extended to every new option.
- And one gap closed on the way: an embed() with neither lang nor system fell silently into English — on a French site, a French-speaking visitor got an English widget. When there is nothing to guess from in the config, the page now decides (html lang, then the browser’s language). The system prompt’s verdict remains sovereign when there is one: instructions must follow the declared behaviour’s language, not the hosting page’s.
The embeddable assistant stopped answering “I do not have that information” to everything outside its notes — measured from 3/11 to 33/33. And four things reported the same day: a storage gauge that contradicted its own figure, a button nobody could interpret, downloaded models buried at the bottom of the list, and documentation pages that scrolled into 583 px of nothing.
SDK 0.1.4: the widget can hold a conversation again
- Reported with a transcript: after one good sourced answer, every following message got the same wall. “are you for real?” → I do not have that information. “are you ok?” → I do not have that information. “hello”, “PLEASE”, “HELP ME”, “I DIE” — all of them. The cause: as soon as the ranking retained no passage, the injected instruction was a refusal, and only a whitelist of greetings escaped it. That refusal is right for “Who won the 1998 World Cup?” — it is the promise of the product — and absurd for “are you ok?”. Worse, a 230M model copies what it has just written: three refusals in the history and the conversation never recovered.
- The two situations are now told apart in code, not left to the model: a message asking for information (at least two content words and an interrogative form) gets the honest refusal; anything else gets one short instruction to reply briefly and kindly. The sorting is deterministic and was checked word by word against the reported transcript — at this model size, an instruction saying “refuse if it is a factual question, otherwise chat” is two rules, and two rules blur into each other.
- Two more things were needed, and both come from a lesson this engine already knew. A pinned example shows the missing behaviour — without it the model had only ONE way of answering without notes under its eyes, the refusal, and copied it. And a deterministic net catches what the instruction does not: on a message we ourselves classified as conversational, a refusal to inform is certainly wrong, so it never reaches the visitor. Same mechanism as the reflex re-greeting stripped in the engine: at 230M, what an instruction fails to get, a deterministic pass gets.
- Measured, with a new benchmark that replays the reported transcript in order on a single conversation — the broken-record effect only shows in sequence. 3/11 before. 25/33 over three rounds with the instruction alone (the record was broken: “hello”, “are you ok?”, “ALLO?” at 3/3, and only one-word interjections and the turn right after a legitimate refusal still fell through). 33/33 with the net, of which only 4 catches: the model gets 29 of them on its own. The benchmark counts the catches separately on purpose — a harness that hides its own safety net measures nothing. The RAG benchmark stays at 12/12 in both languages: the added example degraded nothing.
- A guardrail on the build, after it bit: a plain `npm run build:sdk` silently rewrote `public/sdk-0.1.3.js`, one hour after that version was published on npm. That file must stay byte-identical to the npm package of the same number — it is what an integrator pins so their widget never changes under their feet. The build now refuses to overwrite a frozen file whose content would differ, and says to bump the version instead. It had already happened once, on 0.1.1, which had to be restored from the published tarball.
Four reports, four fixes
- “I deleted everything and the bar stays full, 2.3 GB / 2 GB.” The storage gauge was drawn from the browser’s total for the origin — its own HTTP cache included, which no page can purge — so it contradicted the figure written right above it, “Brimkern data”, which did fall to zero. The bar now has two segments: our share in accent, and in grey what the browser counts on top, with a figure for each as soon as the gap exceeds 4 MB. Delete everything and the accent segment disappears: what fills the bar is visibly labelled as not ours.
- “I do not understand the Clean now button.” It meant nothing on its own: it applies the rule of the dropdown next to it, immediately instead of waiting. That is now written, with the chosen delay inside the sentence. And a click that found nothing to delete used to fall back on the previous report — so the panel re-displayed an old clean as if it had just happened, or displayed nothing at all. Either way the button looked broken. It now says: nothing to clean.
- Model selection: what is already downloaded now comes first. The grid was shown in catalogue order, so a model already on disk — the only one that starts with no network and no wait — could sit in seventh place under six models to download. The sort is stable: inside each group the catalogue order is kept, and nothing moves at all until at least one model is cached.
- The documentation scrolled 583 px past the end of its content. Highlighting the current chapter requires a section title to cross a reading line in the upper third of the screen, and the last sections never reach it — there is no scrolling left to give them. The fix had been a shim reserving exactly the missing space; it worked, and it cost that emptiness. End-of-document is now handled inside the rule, and a table-of-contents click pins its chapter for as long as the scroll takes to settle: on a page that scrolls only 181 px, clicking a chapter used to highlight another one. Four assertions per page now guard it.
embed() now returns a handle, so a widget can be unmounted, driven and observed — and an answer says which of your notes produced it. The SDK demo is bilingual too, with English as its default, and the widget’s own labels follow the page instead of always speaking French. Elsewhere, a GGUF this engine cannot read now says so as it loads.
SDK 0.1.3: a widget you can drive, unmount and observe
- embed() returned nothing, so a mounted widget was mounted forever: in an app with client-side routing it survived every route change, and a second embed() stacked a second launcher on the page — a React effect’s cleanup had nothing to call. It now returns a handle: open/close/toggle, ask() to send a question as if the visitor had typed it, destroy() to remove the DOM and cancel any generation in flight. The engine stays loaded, because the weights are shared by the page: unmounting a widget never makes the next one download the model again.
- The widget was a black box: an integrator could not log a conversation, measure engagement, or even learn that a visitor’s browser had no WebGPU — that failure died inside a chat bubble, which is the single most useful thing the widget could tell you. Both the handle and a session now expose on(event, callback), returning its own unsubscribe function: ready, progress (stable phase key plus bytes), open, close, message (the question as it is sent, the answer when complete), error — which carries code "no-webgpu" for that one failure a visitor can trigger without doing anything wrong. A listener that throws is caught and logged: the host’s analytics code is not our code, and it cannot interrupt a generation.
- A session now preloads before its first turn instead of loading implicitly inside it. That first ask() used to download 149 MB with no callback able to say so — progress had nowhere to come from. Same result, same idempotent load shared by the page; it is simply observable now.
- An answer can now name its sources. The passages chosen from your notes were selected, injected into the prompt and then thrown away: when the assistant answered oddly, nothing said whether the wrong note had been picked or the right one misread. session.ask() takes an onSources callback, sessions expose lastSources, the message event carries the same, and showSources: true prints them under each answer in the widget. On a 230M model an answer a visitor can check is worth more than one merely asserted.
- A conversation can be resumed and a knowledge base swapped without losing it. history is accepted at creation and setHistory() replaces it later, so storing widget.history and handing it back gives a visitor their thread after a reload — the Msg type is exported for exactly that. A greeting is then skipped rather than stacked on top: an extra assistant turn in a short window pushes the real context out, and it measurably degrades which notes get picked. Invalid turns are dropped in silence, because what arrives there comes from persistent storage, and failing to mount a widget over a month-old record with one field too many would be a punishment unrelated to the fault. setKnowledge() re-chunks your documents in place: updating a catalogue used to mean destroying the session and creating another one, which meant throwing the conversation away to change a price. Both throw during a generation — the running turn relies on the last history entry to retract itself.
- Two defects found while measuring the above. A page with two widgets showed them in the same colour: the stylesheet is shared and injected once, so the accent interpolated into it belonged to whichever widget mounted first — it now travels on the element, and no integrator value enters CSS text at all. And the model sometimes copied our own closing marker into its reply (“… within 5 working days. --- END OF NOTES”); that marker delimits the knowledge block in the prompt and has no business in a chat bubble.
- Our own demo page had been working around the missing API, which is the best argument that one was needed. To ask a question from the page, its suggestion chips had to guess the widget’s internal classes, click the launcher, wait 80 ms at random and write into the field — and even then it only FILLED it, the visitor still had to press send. It is now widget.ask(text): one line, and the question actually goes. The demo also turns showSources on, so each answer shows the notes it came from.
- Two benchmarks back this. The RAG benchmark gained an English case set — the demo was bilingual but only its French half was measured, and the two languages do not behave alike — and it now reads the sources of each answer, so a failure says which link broke: the wrong note selected, or the right one misread. 12/12 in both languages. A second harness checks the public surface itself in a real browser without downloading a single byte of weights: 33 checks, including that two widgets keep their own colours and that a destroyed widget really leaves the page.
- One behaviour changed: a programmatic session with knowledge documents now samples at 0.25 like the widget, instead of 0.55. The widget was measured — at 0.55, reading a row out of a table went to the wrong column once in three — and the session had quietly kept the higher value. Same prompt, same behaviour on both surfaces; an explicit temperature still wins.
Loading a GGUF: an unsupported type says so, instead of failing later
- A GGML type the engine cannot dequantize used to fall back to “1 element, 4 bytes”, which made every tensor size in the file wrong: the model loaded, then died much further down on an out-of-bounds read with no apparent connection to the cause. Pasting one of the IQ-quantized files that exist in the wild — LFM2.5’s unsloth repo publishes IQ4_XS, IQ4_NL, UD-IQ2_M and UD-IQ3_XXS next to the K-quants — is exactly how you hit it. The parser now names the full GGML enumeration and throws at the one place where the information is still readable: “unsupported GGML type: IQ4_XS”, with the list of the ones it does read.
- This came out of asking whether the widget’s model could go below 149 MB. The answer was measured and it is no — removing one bit per weight loses the task, not a little quality — and the tooling that answered it is kept: a Node-side dequantizer for the K-quants (tested), a per-tensor dtype hook in the .brik converter, and build tiers that can replay the bit allocation of an imatrix-calibrated GGUF tensor by tensor. Whatever the next candidate model is, it gets measured with the same instruments rather than with an intuition.
The same release: the widget speaks your page’s language
- The widget’s own labels were French, in the code: an English site that added the script tag got a French placeholder (“Écris un message…”), a French privacy line and French loading bubbles under English content. They now follow the same lang option as the instructions — declared, or guessed from your system prompt as before — so one option makes the whole widget match the page.
- Loading phases stopped being French sentences and became stable keys (init, download, tokenizer, gpu): the widget renders them in its language, and an integrator who displays preload()’s status can label them in theirs. The one failure a visitor can trigger without doing anything wrong — a browser without WebGPU — now carries a cause code, so it is stated in the page’s language rather than in the engine’s.
- The live demo is bilingual: /sdk-demo serves English (the canonical version, as everywhere else on this site) and /fr/sdk-demo serves French, with a switch between the two. Both stores hold the same rules and the same figures in two languages; the French set is unchanged to the character, because it is the one the RAG benchmark scores 6/6.
Image generation is 3.6× faster and leaves beta. Then a same-day audit found what the speed was hiding: a desktop default pointing at a model this engine cannot run, eight ratio×resolution combinations that degraded the image in silence, and GPU reductions that were plain wrong on Intel and Mali chips. All fixed — and the models we publish are now checked against a fingerprint before they load.
Image generation: ×3.6, and out of beta
- The 3×3 convolutions now compute 8 output channels per workgroup, sharing one weight load instead of re-reading it per channel: ×3.05 on a full generation (17.0 s → 5.6 s). The 1×1 convolutions got the same treatment two hours later: 4.80 s, ×3.58 against the starting point.
- Image generation leaves beta: 512 px by default on a capable GPU (the model is trained at 512 — below that it stops composing, it crops), cancellation, save, and a progress bar that shows real work with a time remaining.
- Ratio and resolution are now two separate choices: 7 ratios (1:1 to 9:16) × 5 resolutions. Every label announces the size you will actually get — on a 512-native model, asking for HD renders at 512 then upscales 2× on the GPU (bicubic Catmull-Rom plus edge-aware sharpening), and the option says so instead of promising a native render.
- The full Stable Diffusion VAE decoder was written, quantized to int8 and measured: 3.4× slower than TAESD (25.5 s vs 7.4 s) and memory-bound rather than compute-bound at 512². Shelved as a default, and it is what pointed at the convolutions as the real bottleneck.
What the audit caught
- The desktop default had switched to an SDXL model. Two problems: its weights answered 404, and this engine has no SDXL path at all (one text encoder, no additional conditioning) — so the fallback was about to fetch 10.27 GB of fp32 for a topology it cannot read. Back to SD-Turbo, and the resume-a-conversation guard now resolves its URLs through the same function as the loader, because the two had drifted apart and the guard was letting through exactly the heavy download it existed to prevent.
- Eight of the 42 ratio×resolution combinations produced latents whose sides were not multiples of 8. The UNet halves three times then doubles back: on an odd side, the upward path lands one pixel beside its skip connection, and the concatenation reads past the buffer. WebGPU clamps out-of-range reads, so there was no error and no log — just a silently degraded image. Two of the eight were reachable from the interface. Fixed, plus a test that replays the forward pass dimension by dimension on all 42; it also caught 16:9 “balanced”, which was rendering an exact 3:2.
- The new subgroup-accelerated normalisations assumed a 32-lane subgroup. That holds on Apple, NVIDIA and AMD; on Intel integrated graphics (8 or 16) and Mali (4 or 16), most of the partial sums were silently discarded, so every normalisation in the model came out wrong. They are now sized for the smallest possible subgroup and, like every other kernel here, validated against the reference implementation when the engine loads, with an automatic fallback.
- Mobile had gone back to generating at 512 px — precisely the memory spike that makes the operating system reclaim the GPU mid-generation, which is why the ceiling existed. The ceiling is back, and the selector only offers what the machine can hold.
The models you load are now verified
- Every model Brimkern chooses on its own now carries a SHA-256 fingerprint of its manifest, checked before the first tensor is requested; a mismatch refuses the load rather than running an unknown model under our name. The manifest describes the whole package — architecture, tokenizer (for language models it embeds the vocabulary, 5 to 12 MB, so that is covered too), tensor map, shard sizes. It does not cover the tensor bytes themselves, and we would rather say so than imply more.
- Third-party weights (the SD-Turbo files, the TAESD decoder, the vision model) are pinned to a specific commit instead of a branch. A branch is a moving pointer: whoever controls that repository can change the file under our visitors. We never update those files, so the moving pointer bought nothing and cost that exposure.
- A manifest arriving over the network is now validated before its numbers size a single GPU buffer: plausible bounds on every field, and above all every tensor must fit inside its own shard, so no planned read can leave the declared data area. A tokenizer identifier is restricted to the “author/repository” shape — it becomes a network request, and a forged one could have loaded a stranger’s tokenizer.
Vision: 4.6× faster, reading just as well
- The image sent to the visual encoder had been enlarged from 448 to 896 px “for text and interfaces”, without a measurement. The cost is quadratic. Measured on a test card of four lines from 72 down to 13 px: 448 px reads all four lines in 11.4 s (256 visual tokens); 896 px reads the same four in 52.8 s (729 tokens). ×4.6 on the answer for zero extra line — 896 gets the capitalisation more faithfully, which does not buy 40 seconds. Back to 448, and the benchmark ships in the repo so the next change to that number comes with figures.
SDK 0.1.2: the widget answers from your notes
- Knowledge-base selection is 9.4× faster (1.58 → 0.17 ms per question on 120 passages): the per-passage word analysis was being redone for every passage and every query term. It is now computed once per passage. The rewrite was proved equivalent, not assumed: 48 000 score comparisons against the previous implementation, zero divergence.
- The widget was receiving the right note and answering beside it: asked “I wear a 42, what is that in cm?” with “EU 42: 27.0 cm” in front of it, it replied “42 is a size in cm”, and it answered two French questions in English. A browser benchmark now scores six cases on the demo shop (table lookup, two numbers in one note, out-of-scope refusal, small talk): it went from 2/6 to 6/6, stable over ten separate sessions. Four things were wrong, all measured one at a time. The instruction wrapping the notes was always in English, so the model followed its language rather than the question’s. The demonstration turns were written by hand and had drifted from the real ones — they showed notes without the instruction line that precedes them, and at this model size an undemonstrated shape is an unfollowed shape; they are now built by the same function as the real turns, so they cannot drift again. Two operations were never demonstrated at all: reading one row out of a table, and picking the right number when a note holds two. And generation ran at temperature 0.55 — fine for small talk, wrong for copying a figure out of a note; knowledge sessions now run at 0.25.
- A note that merely shares a word with the question no longer gets injected. Asked where free shipping starts, the selection was returning the shipping note (score 0.695) and the returns note (0.228, for the single word “free”) — and the model answered from the second one. A passage scoring a third of the best is noise whatever the absolute threshold, so selection now keeps only what is comparable to its best match.
- The instruction language can now be declared outright (lang: 'fr' | 'en') instead of being guessed from the system prompt. The guess remains as a fallback, with tighter word boundaries — “aide” used to match inside English words.
- A pinned SDK version is now frozen for real. sdk-0.1.0.js on the site had drifted from brimkern@0.1.0 on npm — same name, different bytes — and the freshness test was in fact demanding that pinned files be rebuilt, which is the opposite of what pinning promises. From 0.1.1 on, the self-hosted file and the npm package are byte-identical — verified against the published tarball at each bump.
- Also: dragging a .gguf or .brik onto the model browser works again, and models can be unloaded from their card again — including image, video and vision, which is where a gigabyte of VRAM actually sits. Thumbnails keep their aspect ratio instead of being squashed square. An attached image is stored once instead of three times (about 750 kB down to 40 kB each). Opening the app with the sidebar closed made React discard the prerendered page and re-render everything client-side; that is fixed. And the documentation menu highlights the right section on short pages.
The machine’s compute ceiling is finally measured, and it showed our matrix multiplies were capped by their own inner loop, not by the hardware. Rewritten, they read your prompt up to 1.5× faster at the kernel level on 7B-class models.
Prompt reading: the multiplies catch up
- A new benchmark measures what this GPU can actually compute: 2 825 GFLOP/s. Our prompt-phase matrix multiplies were running at 973, and a mock-up of their inner loop capped at exactly that number: the bottleneck was the kernel’s structure (too many shared-memory reads per multiply), not the machine. The benchmark ships in the repo (scripts/e2e/flops.mjs).
- The int8 and int4 kernels were rewritten with wider register blocks (each thread now computes 4×8 outputs, fed by vectorized reads). Measured kernel-isolated on the exact shapes of a 7B: ×1.4 to ×1.5, lifting the theoretical prompt-reading ceiling of that model from 74 to 113 tok/s. As always: validated against the CPU reference on every load, with a kill-switch (?qshared2=0) and an automatic fallback.
- Measured end to end on a small model (Qwen3 0.6B, ~480-token prompt): prompt reading goes from 506 to 556 tok/s (×1.10). On small models the multiplies are only a quarter of the prompt phase since the August 16 attention fix, so the biggest wins land on the largest models, where they dominate.
- Confirmed the same evening on a mid-size model (Qwen3 4B, in the full app): the prompt-phase multiplies drop from 3.3 to 2.4 ms per dispatch. ×1.48 at equal shapes, right where the kernel benchmark predicted. At that size they are 8× the cost of attention during prompt reading, which is exactly why this was the right kernel to rewrite.
Images generate three times faster
- The convolutions, which carry two thirds of a generation, were computing one output channel per group of GPU threads. The input patch was therefore re-read from memory once per output channel, and each thread did nine multiplications for eighteen reads. Computing eight channels at once from the same patch turns that into seventy-two multiplications for the same eighteen reads.
- The same treatment was then applied to the 1×1 convolutions of the residual shortcuts, which the first fix had promoted to second place in the budget. Measured end to end at 512px: a generation drops from 17.2 to 4.8 seconds, ×3.58. Generating at full native resolution now costs what half resolution used to. Video generation shares the same network and benefits too. The structure was chosen by measurement before a line of the kernel was written: one channel per group reached 300 GFLOP/s, four reached ×2.3, eight ×2.8, and beyond that the returns fade.
Image generation grows up
- Images are now generated at 512px by default on a capable GPU. This is not a comfort setting: the model is trained at 512, and below that it stops composing. On the same prompt and the same seed, 256px returned a cropped fragment of a face where 512px returned the whole portrait. Phones and smaller GPUs stay at 256px, where memory decides rather than taste.
- A running generation can be cancelled. The stop button only interrupted the language model before, so an image or a clip had to be waited out to the end. It now stops at the next block, and frees the GPU memory it was holding.
- Generated images and clips can be saved. A clip in particular only lived for the session, so minutes of computation disappeared when the tab closed.
Waiting for a generation, made legible
- A generation now shows a running clock, a progress bar that tracks the actual computation, and an estimate of the time remaining, instead of a line of text that changed every few seconds. The bar comes from the pipeline itself, so it advances at the real pace rather than pretending.
- Video clips are still shown as 0:00 by no more: browsers record WebM as a live stream whose length never reaches the file header, and the player now measures it on load. Clip length is also yours to choose, from 8 to 32 frames, with the compute cost stated next to each option.
- The separate video lab page is gone: video generation lives in the chat now, next to the other modalities, so the lab was a second door to the same room.
Image and video generation, 1.67× faster
- A new benchmark profiles a whole generation kernel by kernel, and it found the culprit immediately: the convolutions ate 73.8 % of the GPU, and one of them alone (the 3×3 convolution on quantized weights) took 70 % at 35 ms per call. The plain-precision path had a fast tiled version of that convolution; the quantized path, the one the app actually runs, never got one.
- Written, it reads each input pixel once per workgroup instead of nine times, and unpacks each weight once instead of 256 times. Measured on a 256px image: the whole generation drops from 5.0 to 3.0 seconds (×1.67), and that convolution from 35 to 19 ms. Video generation shares the same network, so it benefits too. As always: checked against the CPU reference at every load, with a kill-switch (?convtq=0) and an automatic fallback.
- A generated clip showed a duration of 0:00 in the player: browsers record WebM as a live stream whose length is never written into the file header. The player now measures it on load, so the timeline and the seek bar work.
Video generation joins the chat (beta)
- The video lab becomes a real mode: pick “Generate in the chat” on the video card (desktop), describe a scene, and a short looping clip is generated on your GPU. With step-by-step progress, since a clip takes minutes, not seconds. Your one-line prompt is first expanded by a small local language model into a fuller visual direction (short prompts make static clips).
- Reopening a video conversation shows each clip’s first frame rather than silently losing it: a generated clip lives only for the session (unlike images, regenerating one costs minutes, so there is no click-to-reveal).
Waiting, and reading, in your own language
- Loading an image or video pipeline downloads up to 1.5 GB, and showed a single line of text while it happened. It now has the same progress bar and time-remaining estimate as the language models, so you can tell a slow download from a stuck one.
- The image, video and vision pipelines wrote their loading steps in French only, so an English page could suddenly display “Téléchargement du module motion…”. Every step of those pipelines now follows the language of the page.
- The changelog and the WebLLM comparison now sit inside the documentation shell, with the same side menu, and the changelog gains one anchor per release, so a given version can be linked to directly.
The SDK lands on npm, and the docs get a spine
- The embeddable SDK is now published as the brimkern package on npm. TypeScript types included, safe to import on a server (Next.js, Remix, Astro), same API as the script tag: embed, createSession, generate, preload, status.
- A full API reference page ships with it at /docs/sdk: every option of every call, documented from the published type definitions.
- The documentation now has a sticky side menu: the doc pages, plus the sections of the current page highlighted as you scroll. On narrow screens it folds back into the pill row.
Answers arrive about 40 % faster: the normalization step of every layer was running on a single GPU thread. Reasoning models no longer get stuck on “Thinking…”, the storage panel stops hoarding ranges that will never be read again, and a measured comparison with WebLLM is now online.
Faster answers
- Every layer normalizes its values twice per token, and that step was written one row per thread: fine when reading your prompt (hundreds of rows at once), wasteful when writing the answer (one row, so 63 threads out of 64 idle). Rewritten to split the row across the whole workgroup: decoding goes from 36.0 to 49.5 tok/s (×1.38) on a Qwen3 0.6B, prompt reading unchanged.
- The measurement that found it: a per-pass GPU profiler (?gpuprofile=1) added the same day, which showed normalization eating 51.9 % of decode time. Twice the cost of the matrix multiplies it feeds.
- On a 7B the same fix is worth ×1.27 (8.1 → 10.2 tok/s): the bigger the model, the more the matrix multiplies dominate, so normalization weighs less. That 7B is now in the catalogue: it is the model our published figures are measured on, and you can run it yourself.
Reasoning models: no more dead end
- When a model stopped in the middle of its reasoning, the bubble stayed on “Thinking…” forever: no answer, no explanation, no way out. That state is now shown as a collapsible “Reasoning (interrupted)” block, the reply is marked as cut off, and a Continue button picks it back up.
- Past reasoning is no longer sent back to the model on the next turn (the official Qwen3/R1 templates drop it too): the second-turn prompt shrank from ~240 to 68 tokens, leaving more room for the conversation itself.
Storage that stops growing for nothing
- A model could occupy 239 MB of cache for a 149 MB file: pieces left behind by an older download plan, never read again but still counted against the browser quota that decides whether a model can be kept at all. They are now cleaned up automatically: only pieces fully contained in a larger one, so nothing you already have is lost.
- Models whose download link carries a query string (a common Hugging Face form) were invisible to the storage panel, to “delete this model”, and to automatic cleanup. They are recognized again.
Llama, Mistral and SmolLM3 load differently
- These three families order two of their attention matrices differently from the rest. Until now we rewrote those matrices at load time to match our kernel; the kernel now handles their convention directly. Llama 3.2, Ministral 3 and SmolLM3 were all re-checked and answer correctly either way (?ropenorm=0 restores the old path).
- Consequence: these models can now be packaged as .brik. That rewrite was impossible on a quantized layout, which is why converting a Llama GGUF to .brik was refused.
The site
- A measured comparison with WebLLM at /vs-webllm: same GPU, same 7B int4 model, prefill and decode side by side. Including where WebLLM is ahead.
- Accessibility: the four secondary pages carried contrast violations in dark mode that no audit had ever covered (the red used for solid buttons was also being used for small text). Eight pages × two themes now pass with zero violations.
- On the French home page, the terminal panel stayed permanently faded because longer French sentences pushed it below the fold, where its scroll-driven appearance never completed. It now fades in with the rest of the hero.
- The layer diagram now draws itself piece by piece as you scroll, and the measured figures count up when they come into view.
- The SDK demo asked for “a short story” under a 100-token budget and stopped mid-sentence. It now asks for three sentences, and has the budget for them.
Any single-file GGUF from Hugging Face now runs here: paste an author/model and go. Decoding is 4× faster on large models, the first reply no longer costs ten seconds, the Llama family answers correctly again, and the site finally has a front door separate from the app.
A front door, and the app on its own address
- The home page is now a real landing page explaining what the engine does; the chat lives at /chat. Links published earlier (?model=…) still land in the app.
- A documentation hub at /docs gathers everything: loading a model, instant test links, the .brik format and its converter, the SDK, storage, diagnostics.
- English is now the canonical version of the site (French at /fr), each language on its own indexable URL.
Any model from the Hub, in one paste
- Paste author/model, a Hugging Face URL, or a direct .gguf / .brik link: the best quantization is picked for you and the tokenizer is read from the file itself. Nothing to configure.
- Large GGUFs stream by HTTP range instead of downloading whole: a 4.7 GB model reloads from cache in 15.8 s.
Speed
- Decoding was reusing a kernel built for many tokens at once, leaving seven threads out of eight idle. A dedicated one: ×4.2 on 7B shapes (3.4 → 14.4 tok/s), ×2.4 on a 0.5B.
- The first message of a session paid for moving the weights into VRAM (10.9 s on a 7B). A throwaway warm-up pass moves that cost off your first prompt: 1.1 s.
- Prefill GEMMs are tiled and register-blocked in all three precisions: ×2–2.7 at kernel level, and ~1 TFLOP/s sustained on 7B shapes.
The Llama family answers correctly again
- Llama 3.2 produced fluent nonsense. Cause: a load-time optimization (one HTTP range per layer) filled the weight cache directly, bypassing the row fix these models need on their Q/K matrices, so the chat path read mis-ordered weights. Fixed; Llama now answers correctly, and a CPU reference validates the engine layer by layer.
Everyday things
- Models unused for 30 days are cleaned up automatically (adjustable, or off). The Storage panel now groups by model instead of listing hundreds of HTTP ranges.
- Web search no longer fires on small talk, a truncated reply says so and offers Continue, and reasoning blocks are hidden by default.
- Accessibility: 7 violations → 0 (contrast, form labels, landmarks, heading order), light and dark.
Brimkern goes open source (MIT), and becomes embeddable: one <script> tag puts a local AI on your own site. The ultra-light chat no longer freezes, and video generation gets a resident engine plus a real WebM export.
Open source: the code is public
- The entire engine is now on GitHub under the MIT license: WGSL kernels, the .brik format, the loaders, the app. New home: brimkern.romainkhanoyan.fr.
- A product README with screenshots, and proper SEO plumbing (robots, sitemap, canonical domain).
Embeddable SDK (v0): your site, your visitors’ GPU
- One <script src="/sdk.js"> + Brimkern.embed({ system: … }) mounts a chat widget that runs a .brik model entirely on the visitor’s GPU: zero server, zero per-token cost, private, offline after the first load. Live example on /sdk-demo.html.
- Configurable with a plain object: model (a hosted .brik URL), system prompt, title, greeting, accent color, token budget. The model only downloads when the visitor engages the widget: your PageSpeed is untouched.
- It reuses the app’s fast path (resident GPU decode) and never interprets model output as HTML: plain text only, styles scoped to the widget.
LFM2.5 chat unfrozen, and 2.3× faster
- Switching to LFM2.5 mid-conversation could freeze the tab: every token triggered ~100 GPU round-trips, and replaying the whole history multiplied them by thousands. The forward pass is now fully GPU-resident: one submission per token, one for the whole prefill.
- Verified token-identical to the previous path before shipping, with an automatic fallback and a ?lfm2resident=0 switch. Measured: 13.5 → 31 tok/s on the same machine.
Video (lab): resident engine, prompt enrichment, WebM export
- The temporal (motion) modules now run entirely on the GPU in a single submission: 5× faster per module, with a CPU-verified fallback and ?videoresident=0.
- Your short prompt is enriched by the local LFM2.5 model into a proper cinematic description before generation: better motion, still 100% on-device.
- The result exports as a looping WebM clip of bounded duration (~10 s) instead of raw frames.
RWKV-7 joins the catalog: linear attention, constant memory
- RWKV-7 G1 0.1B is loadable from the model browser: a 100% recurrent architecture where a fixed ~1 MB state replaces the KV cache. Memory does not grow with the conversation. 128 MB, Apache-2.0, embedded World tokenizer, naive-but-honest replies (it is a 0.1B).
Mobile & housekeeping
- The full model browser is now reachable on mobile with a “Change model” action: Qwen 3 0.6B and friends are one tap away, no longer desktop-only.
- The storage gauge now shows the browser’s real quota (it depends on your free disk space) instead of an optimistic estimate, the misleading first-visit splash is gone, and the “GPU engine” card is more compact.
A second engine is born: linear-attention and hybrid models run in your browser. A 149 MB model that chats in French, a live demo on /local-ai, and the first video ever generated by Brimkern, entirely on your GPU.
LFM2.5 230M: the ultra-light that actually chats (mobile too)
- New engine path for hybrid architectures (short-convolution + attention): LFM2.5-230M runs end-to-end on our WGSL kernels. 149 MB, replies in clean French, ~24 tok/s.
- Available in the main chat as a preset, and on mobile, where you now pick between LFM2.5 (149 MB, recommended) and Qwen 0.5B (378 MB).
- Every kernel validated against a CPU oracle, token-for-token vs llama.cpp before shipping.
Live demo on /local-ai: classify, extract & chat
- Sentiment (12/12 on our benches), email extraction with an anti-hallucination guard (an email absent from your text is never invented), and a small French-speaking chat: all in your browser, cached after the first 149 MB download.
- Constrained classification under the hood: the model can only answer within the allowed label set. The technique that makes tiny models reliable.
First video generated in the browser (lab)
- 16 temporally-coherent frames (a fox walking through snow) in ~3.5 minutes, 100% local: AnimateDiff-Lightning motion modules grafted onto our existing image pipeline. Zero new GPU kernels needed.
- Still lab-only (dev bench): the product UI, thermal pacing and WebM export are next.
Fixes & housekeeping
- Reopening a conversation that contains images no longer auto-downloads the image model: only text models auto-load, your saved images display as-is.
- “Cached” badge renamed “Downloaded” with a distinct solid style: no more confusion with the “Runs well” GPU verdict.
Brimkern learns to see: show it a photo and ask questions. Entirely on your GPU. Refining an image now starts from its real pixels, and the last deferred kernel lands.
Vision (beta, desktop): image + text → text
- Qwen2-VL 2B runs fully in the browser: a 675M-parameter vision encoder (32 transformer layers, 2D rotary positions) reads your image as patches, a merger projects them into the language model, and the LLM answers your questions about it. Both weight files stream from Hugging Face and stay cached (~2.3 GB, desktop GPUs only).
- Under the hood: two new position kernels (2D RoPE for the vision tower, M-RoPE for the language model). The second one touches the chat hot path, so it only ever dispatches for this architecture and self-tests at load with a fallback that disables vision, never text.
- In the model browser, the Qwen2-VL card is now loadable (desktop). Attach an image with the 📎 button and ask away: multi-turn works, everything stays local.
Two new brains in the catalog: Qwen 3, and Llama is back
- Qwen 3 (4B and 0.6B): the next generation, clearly stronger than Qwen 2.5 at equal size, with native step-by-step reasoning. The thinking budget selector applies to it. Its QK-Norm architecture is self-tested at load like every kernel change.
- Llama 3.2 is repaired and back in the catalog: llama.cpp permutes attention weights at GGUF conversion in a way our kernels didn’t expect. They are now un-permuted at load (all quantizations), and the llama3 long-context frequency scaling is applied. Measured: 178 t/s prefill, 19.5 t/s generation.
Real img2img: refine from the pixels
- “Refine this image” used to replay the same starting noise; it now encodes the displayed image back into latent space (tiny 5 MB VAE encoder, fetched on first use), re-noises it partially and regenerates: composition is preserved from the actual pixels. Tune with ?strength= (default 0.55).
- A refined image depends on its source pixels, so it can’t be regenerated from prompt+seed like the others: it is saved whole with the conversation.
Under the hood
- The last deferred kernel is in: full (non-causal) attention now gives each (token, head) a 64-lane workgroup with online softmax. One pass over keys instead of two. Self-tested at real UNet shapes at load, silent fallback, ?attnfullwg=0 to force the old kernel.
- Clearing the conversation history now really clears the screen (fresh chat), and the app no longer auto-reopens a conversation whose model isn’t cached: you land on a fresh home instead of a dead chat.
Long conversations get their speed back (up to ×26 on attention), the GPU can disconnect without killing the app, the model loads itself, and the interface adopts its print identity for good.
Attention rebuilt for decoding: the end of 1 t/s on long context
- The attention kernel used a single GPU thread per head: 14 threads total while decoding. On hardware built for thousands. Past ~1,000 tokens of context it became the wall (~680 ms per token). New kernels give each head a full 64-lane workgroup with online softmax: ×14 to ×26 measured, identical results to 1e-7.
- Belt and braces: at load time the engine self-tests the new kernels at real-world shapes; a GPU driver that miscompiles them falls back silently to the classic kernels (slower on long context, correct everywhere). ?attndecode=0 forces the fallback for diagnosis.
- On mobile, answers are now deliberately concise (the phone shouldn’t heat up for a minute per reply), and the screen stays awake while the model works.
A GPU crash is no longer the end
- When the system reclaims the GPU (long conversation, backgrounded tab, thermal pressure), the app used to keep running on a dead device: “ready” status, sends allowed, every compute failing. It now detects the loss, keeps your conversation, and offers real exits: reload the model, inspect/clear storage.
- WebGPU detection retries before giving up, and “Unsupported” now explains the #1 cause: hardware acceleration disabled in the browser. With the exact setting to flip.
The model loads itself
- Reopening the app resumes your last conversation AND its model when it’s fully cached (streamed BRIKs included: desktop and mobile alike, zero network). Partially downloaded? The background prefetch finishes the job: with real progress shown on the first-visit splash. Then loads the model on its own.
- Fixed along the way: the background prefetch could silently die before ever starting, and an auto-load at startup mistook the phone for a desktop (loading f16 instead of the mixed format).
Image generation slims down, and lands on mobile
- The image pipeline downloaded 2.4 GB of fp16 weights, then quantized them to int8 on your GPU at every load. The weights now ship pre-quantized and range-streamed (resumable, cached): 1.28 GB on desktop, identical output. Verified numerically AND visually against the old path.
- New on mobile (beta): “Try image generation” loads SDXS-512, a distilled 1-step UNet, in an int4 “light” build with a lightened text encoder. ~445 MB all-in for a native 512px image. Judged side-by-side against the heavy build: near-identical.
“Le Kern”, fully inked
- The display face becomes Fraunces (a printer’s serif: the kern-B mark, titles and splash carry it), a red printer’s rule crowns the app, the caret and text selection turn Kern red, and the welcome screen reads like a type specimen. The last purple remnants and gradients are gone.
- Mobile decluttered: the BRIK-conversion banner is gone (the mixed model is served automatically), and a single starter suggestion leaves room for the model’s welcome message.
The mobile model gets int8 quality at nearly the int4 size (new “mixed” format), downloads itself in the background while you read the home screen, and greets first-time visitors properly.
Mixed quantization: int8 quality, (almost) int4 size
- A tensor-by-tensor A/B study showed WHERE int4 breaks a small model: full-int4 produces nonsense, but keeping just the attention matrices in int8 restores int8-grade quality. The new “mixed” format stores exactly that: int4 body + int8 attention.
- The mobile model is served in mixed format: 377 MB instead of 508 (full int8). For +18 MB over the old int4 file that degraded it. The diagnostics strip shows “mixed int8+int4” honestly.
- The “Mixed” profile is also available in both BRIK converters (recommended for small models: int4 stays for the big ones).
The download disappears into the background
- On mobile, if the model isn’t (fully) cached yet, its download now starts by itself shortly after you arrive: by the time you tap “Load the model”, most of it is already local. Resumable: a closed tab only re-downloads what’s missing.
- Visible and polite: a progress line with a Cancel link, nothing starts if your phone’s Data Saver is on, and any real load takes over instantly.
- First visit on mobile: a short welcome screen (“Preparing your AI space…”) covers the kickoff. Tap to skip, never shown again.
Under the hood
- Prompts are no longer re-tokenized from scratch on every message: only the new turn is tokenized (~×90 faster on long histories).
- Generated images no longer freeze the page for ~100 ms after rendering (async PNG encoding).
Mobile finally smooth and reliable: ~4× faster conversations (KV cache reused across turns, GPU-side sampling), int8 by default, and a Settings panel.
Faster conversations: especially on mobile
- The attention (KV) cache is now reused from one message to the next: only your new message is processed, never the full history again. Previously every turn re-read the entire conversation: response time doubled from the 2nd message on; it is now constant.
- Next-token sampling now runs on the GPU (softcap, repetition penalty and top-K fused into the same pass as the forward computation): ~600 KB read back per token → 512 bytes. Self-tested at load time, with automatic fallback to the CPU path if the device’s GPU fails the test.
- Leaner generation loop: end-of-text detection on the last few tokens only, and the display refreshes ~8×/s (instead of a full re-detokenization and re-render on EVERY token, which choked phones). Only the bubble being typed re-renders.
- Measured on a phone (Qwen 0.5B): full response in ~10 s instead of 20–38 s, generation at ~5 t/s instead of 2.7.
Mobile quality: int8 by default
- Small models (≤ ~1.2B parameters) now load in int8 on mobile instead of int4: int4 severely degraded a 0.5B (nonsensical answers, repetition loops) while it easily fits in int8. int4 stays reserved for large models that would not fit otherwise.
- And the attention cache stays in f32 by default (faster): its int8 variant. Which adds work on every token: is only enabled where its VRAM savings really matter (large models, int4).
- New diagnostics strip under every response: actual precision, KV cache format, sampling path and context reuse, so you can see what actually ran, even without a console. And it tells the truth: when the model file is more quantized than the requested precision, it shows e.g. “int8 (source int4)”.
- Persistent storage is now requested from the browser: the cached streamed model should no longer be evicted between sessions on mobile.
Image pipeline: even leaner
- The text encoder (CLIP) joins the UNet: int8-quantized, GPU-resident, executed in a single submission. No more ~280 round trips and ~500 MB of weights re-uploaded to the GPU for every image.
- The image model now loads without freezing: fp16 weight conversion happens on the GPU (the tab used to lock up for ~10 s on every load).
- GPU memory is released after each image (the scratch buffer used to hold on to hundreds of MB), and the image pipeline is fully freed (~1 GB) when a text model is loaded.
- Faster 512px decode: the dominant 3×3 convolutions now run as a tiled kernel with shared memory (each pixel read once instead of 9×, each weight once instead of 256×), and group normalization uses 4× more threads. Both self-tested at load with automatic fallback.
Web & tools (MCP): first steps, fully transparent
- Optional web search (OFF by default): the model draws on Wikipedia excerpts and cites its sources. Only your question is sent: never the conversation. Flagged under every affected response and in the input bar.
- Local calculator (ON by default, no network): arithmetic in your messages is evaluated exactly on-device and the result handed to the model. Small models systematically get arithmetic wrong. The model also knows today’s date.
- Pasted-link reading (OFF by default): paste a URL and the model reads the page (via the r.jina.ai reader. Noted on the option).
Settings & fixes
- New “Settings” panel in the sidebar: GPU power (Eco / Balanced / Max. Keeps heat in check during image generation) and the Web & tools options above.
- Fully bilingual interface: the FR/EN toggle now covers the entire application (panels, errors, loading steps, tooltips…). Including this changelog.
- Snappier streamed model loading: a layer’s weights are fetched in a single request instead of 9–12 (~25 requests instead of ~220 for a whole model).
- The loading screen is now a real log: opaque background and a list of the actual steps (download, quantization, validation) with their progress.
- Auto-scroll respects your reading: if you scroll up while the AI is writing, you are never yanked back to the bottom. It resumes once you scroll back down.
- Fixed: loading a text model on top of image mode did not switch over. Messages were still sent to image generation (and the image pipeline stayed in memory).
- Diagnostic switches in the URL (?gputopk=0, ?kvreuse=0) to isolate device-specific issues.
100% in-browser image generation (SD-Turbo on hand-written WGSL), int8-quantized GPU-resident UNet, thermal throttling, and a new visual identity.
Image generation: Stable Diffusion Turbo, 100% local
- New image mode in the chat: describe an image and it is generated entirely in the browser. CLIP (prompt encoding), UNet (denoising) and the TAESD decoder all run on our own WGSL kernels, with no server involved.
- Prompt fidelity fixed against the diffusers reference: tokenizer padding (“!”, not <|endoftext|>. With no attention mask, all 77 positions count) and CLIP’s final LayerNorm now applied. Result: images that actually match the request.
- Per-generation quality selector above the input area: 128px (fast), 256px (recommended) or 512px (SD-Turbo’s native resolution).
- Lightweight conversations: only a blurred thumbnail + the prompt + the seed are persisted; “click to reveal” regenerates the exact same image (deterministic generation) without ever storing the pixels.
- Weights (UNet + CLIP fp16, TAESD) are cached by the browser on first load → no download on subsequent runs.
Performance & thermals: resident int8 + throttling
- UNet quantized to int8 (BRIK8) directly on the GPU at load time: ~0.9 GB of VRAM instead of ~3.4 GB in f32, and no more weight re-uploads per image (previously: ~3.4 GB re-transferred per generation). New int8 conv2d kernel with fused dequantization, covered by the self-test.
- End-to-end GPU-resident execution: activations stay on the GPU across all blocks (UNet and decoder), with a single CPU readback per denoising step → far fewer round trips, more efficient generation.
- Built-in thermal throttling: the pipeline measures the GPU time actually consumed and inserts proportional pauses (~60% target load). Generation smooths out its power draw instead of heating up the machine in one continuous burst.
- Detailed progress during generation (current denoising step + UNet block).
New identity : “Le Kern”
- Goodbye purple and gradient logo: enter a paper / ink / printer’s-red identity, a nod to the typographic kerning in brimKERN. Matching “ink” dark mode.
- Logo redrawn as flat SVG (a massive B notched by a diagonal kern): crisp at every size, follows the light/dark theme, and replaces 700 KB of PNG.
- Favicon and share card (OpenGraph) updated to the new identity.
Cleanup
- Removed the “preview” placeholder image generator, the Llama 2/3 prompt templates unreachable since Llama was pulled, and unused assets: lighter bundle.
Gemma 2 fully working, automatic tokenizer matching, conversation resume, prefill performance (tiled matmul), Skills, a richer storage panel and a bilingual FR/EN interface.
Gemma 2: fully working
- Coherent generation on Gemma 2 (2B): fixed a numeric overflow in the GELU activation (which produced NaNs) and the double application of the (1+w) RMSNorm. The cause of the garbled text.
- Automatic tokenizer + architecture matching from the GGUF file: loading a model can no longer mistakenly use another model’s tokenizer (a vocabulary mismatch produced gibberish even though the math was correct).
- Reusable foundation for adding more local model families (next target: Microsoft Phi-3.5-mini).
Model selection & storage
- Full resume on open: the last conversation is reloaded and, if the model it used is already cached, it is reloaded automatically (no network) → you land straight back in a ready-to-chat session.
- New wide-format “Browse models” picker with search (name, use case, tag) and a grid: far more readable than the cramped list as the catalog grows.
- Badges on every model: “● cached” (already downloaded locally), “BRIK recommended” (≥ ~1.5B: 2–4× less VRAM + instant reopens), and an indicator on the currently loaded model.
- Storage panel: the active model is highlighted, a “loaded” badge marks the matching BRIK, and a “Delete all” button clears everything (caches + BRIK + history).
- Roadmap: preview of the next architecture to be ported to our kernels (Microsoft Phi-3.5-mini).
Interface: unified model picker
- All model selection now goes through a single fixed-size “Choose a model” window (internal scroll): a Models tab (grid + search + creator/modality per tile) and an Import tab (local file, GGUF URL, .brik stream), with the “Convert to BRIK” checkbox.
- Leaner chat header: shortened model name (tooltip on hover), language/theme toggle moved next to the logo, and dev options (precision/VRAM, KV cache, benchmark) folded into an accordion closed by default.
- Multimodal roadmap preview (Microsoft Phi-3.5, Mistral, Stable Diffusion, Qwen2-VL) as “coming soon” cards.
Performance & conversion
- Tiled q8/q4 matmul for prefill: each invocation computes 4 tokens at once, dequantizing each weight only once → ~4× less weight memory traffic during prompt processing.
- Streaming GGUF → BRIK conversion (shard by shard): peak memory ≈ one layer instead of the whole model → large models convert in the browser without exhausting RAM.
- Llama temporarily removed from the offered architectures (RoPE incompatibility with weights permuted by llama.cpp → incoherent output); it will return with a suitable RoPE mode. Active architectures: Qwen 2/2.5, Gemma 2, DeepSeek-R1 (Qwen distill).
Skills: reusable instructions
- A library of “skills” (personas / system instructions): built-ins + your own skills, persisted locally.
- Multi-select (skills combine), import from a GitHub URL, and a popup reachable via a button to the left of the chat bar.
Storage management
- “Storage” panel: see the space taken by streamed models, downloaded GGUFs, converted BRIKs and chat history. With per-item deletion.
Mobile & bilingual
- Simplified mobile: a single ready-to-use model (streamed Qwen 0.5B BRIK), a download progress bar, and the chat area shows while loading.
- Bilingual interface: French (default) / English, toggle in the header, remembered.
Self-contained BRIK: hosted models, embedded tokenizer, lighter files, and optimized mobile loading.
Self-contained BRIK, hosted & mobile-optimized
- Tokenizer embedded in the .brik → 100% offline loading, no external fetch and no manual tokenizer selection.
- Tied embeddings deduplicated (output = token_embd) → file ~⅓ smaller, with no quality loss.
- Pre-converted Qwen 2.5 0.5B model, hosted and streamed via HTTP Range (low VRAM) → optimized direct loading, on mobile and desktop alike.
- More reliable GGUF → BRIK conversion: very large tensors are processed in slices so they no longer exceed the GPU buffer limit (which silently corrupted weights).
- Adjustable thinking level (off / low / medium / high) for reasoning models; catalog narrowed to fully supported architectures (Qwen, Gemma, DeepSeek).
BRIK v2 format: lighter & streamable
- Web-native BRIK8 (int8) and BRIK4 (int4) quants: dequantization fused into the GPU matmul (weights kept quantized in VRAM), small download AND fast inference.
- Single self-contained .brik file (header + manifest + 16-byte-aligned data) replacing .brik.zip: one file to host/load.
- HTTP Range streaming: only the header (manifest) is fetched first, then each tensor on demand, cached (Cache API) → instant, offline reloads.
- Embeddings stored in int8 (often the largest tensor) → markedly smaller download and VRAM, with near-identical quality.
Longer context: q8 KV cache
- Optional int8 attention (K/V) cache: ~4× less cache VRAM → up to ~4× more context at equal VRAM, near-f16 quality.
- “KV f32 / KV q8” toggle in the precision setting (dequantization fused into attention, with no f32 expansion).
New architectures
- Gemma 2 supported by the optimized kernels: attention + logit softcapping, GELU activation, dual “sandwich” norms, (1+w) RMSNorm, embedding scaling, head_dim ≠ d/heads.
- Kernels made parameterizable (attention scale, activation, norms): a reusable foundation for other model families.
Loading & caching
- Automatic GGUF → BRIK conversion at load time (optional): done once, then the converted .brik is cached (IndexedDB) for instant opens.
- Dedicated conversion page with converted-model cache management (list / delete).
- Weight precision clarified: GGUF in f16/f32, BRIK8/BRIK4 tiers reserved for BRIK models.
Robustness & fixes
- Fixed a GPU dispatch overflow on long prompts (> ~860 tokens) that produced incoherent output: now spread across a 2D grid.
- Token counter + context warning in the composer.
- Large pastes collapse to a “snippet” in the input field (the full text is still sent to the model).
- UI fixes: no more lingering focus ring on click, message bubbles no longer break short words.
First release: an LLM inference engine written from scratch, 100% in the browser.
Custom WebGPU engine
- Hand-written WGSL compute kernels: vectorized matmul (128-bit vec4), RMSNorm, RoPE, causal GQA attention with KV cache, SwiGLU.
- GGUF parsing directly in JavaScript and weight dequantization on the GPU.
- GPU-resident decode path: a token’s entire forward pass chained into a single GPU submission.
- Kernel self-validation at load time (selfValidate): the model only loads if the math checks out.
Performance
- Logit projection cached on the GPU instead of being recomputed for every token.
- Next-token argmax computed on the GPU (a single integer read back per token instead of ~152k logits).
- Buffer pool reused across tokens. Overall result: ~2.5× faster decoding.
Weight precision (switchable)
- f32: full precision (quality reference).
- f16. Half precision: ~1.25× faster, half the VRAM (GPU-dependent).
- int4 (BRIK “q4web” format): on-the-fly dequantization, ¼ of the VRAM → lets you load larger models in the browser.
Models & quantizations
- Models: Qwen 2.5 (0.5B, Coder 1.5B), Llama 3.2 1B, DeepSeek-R1 Distill Qwen 1.5B (<think> reasoning).
- GGUF quantizations supported: Q4_0, Q4_K, Q5_0, Q5_K, Q6_K, Q8_0, F16, F32.
- Import any compatible GGUF (local file or Hugging Face URL).
Interface
- Chat with Markdown rendering (bold, italics, lists, headings) plus syntax highlighting and copy on code blocks.
- Persistent conversation history (IndexedDB), independent of the loaded model.
- Built-in benchmark comparing f32 / f16 / int4, plus a precision selector.
- Collapsible sidebar, mobile-friendly interface.
Privacy
- No data ever sent to a server: the model and all computation run entirely on your GPU, offline once the model is downloaded.
Brimkern: open WebGPU engine, built by Romain Khanoyan. Local AI, WebGPU, on-device engines.