SDK & npm package
The complete API of the brimkern package: a chat widget in one call, or headless sessions and one-shot generation for your own UI. Everything runs on the visitor’s GPU: no server, no API key, nothing leaves the browser. For the guided tour and live demo, see the SDK page.
Install
From npm. TypeScript types included:
npm i brimkernimport { embed, createSession, generate, preload, status, runtime } from 'brimkern'; // types, for TypeScript import type { EmbedConfig, SessionConfig, AskOptions, BrimkernWidget, BrimkernSession, BrimkernEvents, BrimkernEvent, Msg, Source, LoadProgress, ToolSpec, WidgetLabels, } from 'brimkern';
Or as a script tag, with no build step. The IIFE exposes the same API on a global:
<script src="https://brimkern.com/sdk.js"></script> <script> Brimkern.embed({ title: "Ask us anything" }); </script>
embed(config?)
Mounts the chat widget in the page (it waits for the document if called early) and returns a BrimkernWidget handle — see Controlling the widget below. The model downloads only when a visitor actually opens the widget: your page speed is untouched.
systemstringtitlestringgreetingstringaccentstringmodelstringmaxTokensnumberlang'en' | 'fr'examples{ user, assistant }[]knowledge / knowledgeBudgetsee belowworker / workerUrlboolean / stringhistoryMsg[]showSourcesbooleantools('calc' | 'date' | ToolSpec)[]theme / position / width / height / labelssee belowAppearance & labels
The widget follows your page instead of imposing its own look. Everything here is set per widget: two widgets on the same page can differ in theme, corner and colour.
embed({
theme: 'auto', // 'light' (default) | 'dark' | 'auto' — auto follows prefers-color-scheme, live
position: 'bottom-left', // 'bottom-right' (default) | 'bottom-left'
width: 400, height: 600, // px, clamped to 300-480 × 380-720; small screens still cap to the viewport
labels: { // any language, any tone — keys you omit keep the lang defaults
open: 'Chat öffnen',
placeholder: 'Nachricht eingeben…',
note: 'Lokale KI — läuft auf Ihrer GPU.',
phases: { download: 'Modell wird geladen…' },
},
});lang covers English and French out of the box; labels is the door to every other language (open, close, placeholder, note, error, empty, help, sources, mb, phases). Labels are rendered as text, never as markup, and theme takes a keyword, not CSS: nothing an integrator passes here can inject styles or markup into the host page.
Controlling the widget
embed() returns a handle. It is what lets you unmount the widget — which matters in any app with client-side routing, where the widget would otherwise survive every route change and a second embed() would stack a second launcher on the page.
const widget = embed({ title: "Support" }); widget.open(); widget.close(); widget.toggle(); await widget.ask("Do you ship to Canada?"); // as if the visitor had typed it widget.setKnowledge(newDocs); // swap the cards, keep the conversation widget.setHistory(saved); // resume a conversation widget.history // Msg[] widget.el // the panel, for a style tweak widget.destroy(); // removes the DOM, cancels any generation in flight
destroy() leaves the engine loaded: the weights are shared by the page, so unmounting a widget never makes the next one download the model again. In React, the handle is exactly what an effect’s cleanup needs:
useEffect(() => {
const widget = embed({ system: "You are our support agent." });
return () => widget.destroy();
}, []);Calling embed() on the server is harmless: you get an inert handle instead of a crash, so the same code can run on both sides.
Events
Both the widget handle and a session expose on(event, callback), which returns its own unsubscribe function. This is how you log conversations, measure engagement, and — most useful of all — learn that a visitor’s browser has no WebGPU, instead of that failure staying inside a chat bubble.
const off = widget.on('message', ({ role, content, sources }) => { analytics.track('chat', { role, content }); }); off(); // unsubscribe widget.on('progress', (phase, p) => bar.value = p ? p.loaded / p.total : 0); widget.on('ready', () => console.log('model loaded')); widget.on('open', () => {}); widget.on('close', () => {}); widget.on('error', (err) => report(err)); // e.g. no WebGPU on this browser widget.on('tool', ({ name, result }) => {}); // a tool produced a result for this turn
phase is a stable key — init, download, tokenizer, gpu — never a sentence: you label it in your page’s own language. A listener that throws is caught and logged: your analytics can never break the widget. Sessions get ready, progress, message, error and tool (open and close are widget-only), and a session emits progress because ask() now preloads before its first turn: that first call used to download 149 MB with no way to say so.
The one failure a visitor can trigger without doing anything wrong carries a cause code: err.code === "no-webgpu" means this browser cannot run the assistant at all. Treat it as the signal to hide the widget rather than as a bug — status() answers the same question before anything is mounted. Errors on the message path are emitted AND thrown, so ask() keeps its own catch.
createSession(config?)
Headless: a conversation object for your own interface. Same config as embed() minus the visual options — including history to start from a stored conversation — plus temperature.
const session = createSession({ system: "You are a sommelier.", temperature: 0.7 }); const reply = await session.ask("A wine for oysters?", { onToken: (text) => output.textContent += text, // streaming signal: controller.signal, // cancellable }); session.history // the Msg[] so far session.lastSources // the cards behind the last answer session.setHistory(saved) // resume a conversation session.setKnowledge(newDocs) // swap the cards, keep the conversation session.on('message', log) // same events as the widget session.reset() // same config, blank history session.destroy()
setHistory() and setKnowledge() throw if a generation is running: finish or cancel the turn first. Both keep the engine and the weights untouched — swapping a catalogue does not cost a download, and no longer costs the conversation either. ask() also throws if a turn is already running on that session: one conversation, one turn at a time.
temperature defaults to 0.25 when you pass knowledge, and 0.55 otherwise — the same rule as the widget, and a measured one: at 0.55, reading one row out of a table went to the wrong column once in three. An assistant copying a figure out of a note has nothing to gain from sampling wide. Declaring temperature yourself still wins.
generate(options)
One shot: a prompt, a reply, no history kept. Takes the session config plus prompt, onToken, signal and onSources. It takes ONE object: called as generate("question", {…}) it throws a TypeError instead of quietly answering the string "undefined".
const answer = await generate({ system: "Answer in one sentence.", prompt: "Why is the sky blue?", onToken: (text) => process(text), });
Knowledge documents
Give the assistant your content: pages, FAQs, product sheets. Documents are chunked into passages in the browser, and only the 1–3 passages closest to the visitor’s question are given to the model. The ranking is local (lexical): nothing is sent anywhere.
knowledge: [ "Plain strings work.", { title: "Shipping", text: "Free in France from 60 euros." }, ], knowledgeBudget: 800 // max tokens of passages per question
Tools
Give the assistant abilities beyond its notes: arithmetic, today’s date, or your own functions — a stock lookup, an order status, a cart total. The design is deliberate: the model NEVER decides to call a tool (below ~3B parameters, emitted tool calls are hallucinated — measured). Instead, detection is deterministic, your function runs in your page, and the model receives the result as a fact, exactly like a knowledge passage.
embed({
tools: [
'calc', // detects arithmetic in the message and injects the exact result
'date', // the model knows today’s date (stable line in the system prompt)
{
name: 'stock',
match: /stock|disponible/i, // or a predicate: (question) => boolean
run: async (question) => { // runs in YOUR page — sync or async
const n = await api.stockFor(question);
return `${n} in stock`;
},
},
],
});
widget.on('tool', ({ name, result }) => trace(name, result)); // before generation, like onSourcesThe contract protects the visitor: a tool that throws, hangs (10 s cap) or returns nothing is simply absent from the turn — the turn itself never fails because of a tool. Results are capped at 600 characters: a result is a fact, not a report, and the default model has a short window. Nothing here touches the network unless YOUR run function does.
Sources of an answer
You can see which passages fed an answer. Two reasons this matters: a small model does get things wrong, and an answer a visitor can check is worth more than one merely asserted — and when yours answers oddly, this is how you tell a bad passage from a bad reading of a good one.
await session.ask(question, { onSources: (sources) => show(sources), // before generation: the ranking is local and instant }); session.lastSources // [{ title, text, score, doc }] embed({ knowledge: docs, showSources: true }); // in the widget, under each answer widget.on('message', ({ sources }) => trace(sources)); // or without displaying anything
score is the lexical proximity to the question, doc the index of the document in your knowledge array, and the order is the order the passages were given to the model. An empty array is meaningful: no passage matched.
And what happens then depends on the message, because two situations must not be confused. A question asking for information gets the honest answer — “I do not have that information” — and that is the promise of the product. Anything else does not: “Are you ok?”, “PLEASE”, “hello” get a short, friendly reply. An assistant that stonewalls everything outside its notes is one visitors close, and it only takes a few such answers in a row for a small model to keep repeating them.
preload(), status() & runtime()
preload() downloads the engine and the model ahead of the first question: call it on a hover, or on the pricing page before support opens. onProgress receives the phase and, during download, the bytes: enough for a real progress bar.
await preload({ onProgress: (status, p) => { if (p) bar.style.width = (100 * p.loaded / p.total) + "%"; }, }); status() // 'unavailable' (no WebGPU) | 'idle' | 'loading' | 'ready' | 'error'
status() answers synchronously: use it to decide whether to show the widget at all on browsers without WebGPU. runtime() reports where inference actually runs — "worker", "main", or "pending" before anything has started — which is what you check when you passed worker: true and want to know whether the fallback kicked in (a host CSP that forbids blob: makes it fall back to the main thread, silently and on purpose: a widget must never stop working over an execution choice).
One engine per page
The engine is a singleton per model URL: N widgets and N sessions on a page share one WebGPU init and one set of weights in VRAM. Mounting a second widget costs a DOM node, not 149 MB — and destroying one leaves the weights loaded for whoever comes next. This is also why worker and workerUrl only take effect before the first preload or ask: once the backend exists it is shared by the whole page, and a later embed() saying otherwise is ignored with a console warning rather than silently believed.
Versions & CDNs
Pin a version if you would rather the widget did not change under your feet:
https://brimkern.com/sdk-0.6.0.js instead of https://brimkern.com/sdk.js <!-- or from the npm CDNs --> https://unpkg.com/brimkern@0.6.0/dist/brimkern.iife.js https://cdn.jsdelivr.net/npm/brimkern@0.6.0/dist/brimkern.iife.js
Servers, licence, links
Importing the package on a server does nothing until a browser runs it: Next.js, Remix and Astro pass without guards. MIT licence, like the whole engine.
Brimkern: open WebGPU engine, built by Romain Khanoyan. Local AI, WebGPU, on-device engines.