Brimkerndocs

SDK & npm package

The complete API of the brimkern package: a chat widget in one call, or headless sessions and one-shot generation for your own UI. Everything runs on the visitor’s GPU: no server, no API key, nothing leaves the browser. For the guided tour and live demo, see the SDK page.

Install

From npm. TypeScript types included:

npm i brimkern
import { embed, createSession, generate, preload, status, runtime } from 'brimkern';

// types, for TypeScript
import type {
  EmbedConfig, SessionConfig, AskOptions,
  BrimkernWidget, BrimkernSession, BrimkernEvents, BrimkernEvent,
  Msg, Source, LoadProgress, ToolSpec, WidgetLabels,
} from 'brimkern';

Or as a script tag, with no build step. The IIFE exposes the same API on a global:

<script src="https://brimkern.com/sdk.js"></script>
<script>
  Brimkern.embed({ title: "Ask us anything" });
</script>

embed(config?)

Mounts the chat widget in the page (it waits for the document if called early) and returns a BrimkernWidget handle — see Controlling the widget below. The model downloads only when a visitor actually opens the widget: your page speed is untouched.

systemstring
The assistant’s instructions: who it is, what it may say.
titlestring
Widget header text.
greetingstring
First message shown before the visitor types.
accentstring
Accent color (any CSS color).
modelstring
Model override: a direct URL to an LFM2 or RWKV-7 .brik. Defaults to the built-in small model (149 MB). Other architectures (any single-file GGUF) live in the app, not the SDK.
maxTokensnumber
Reply budget, in tokens.
lang'en' | 'fr'
Language of the widget labels (placeholder, status bubbles, errors) and of the instructions given to the model. Guessed from your system prompt when left out — declare it if your prompt is unusual.
examples{ user, assistant }[]
Few-shot examples prepended to the conversation.
knowledge / knowledgeBudgetsee below
Your content, ranked locally: see the dedicated section.
worker / workerUrlboolean / string
Run inference in a Web Worker (keeps your page’s main thread free). workerUrl serves the worker from your own origin if needed.
historyMsg[]
Starting conversation, so a visitor finds their thread again after a reload: store widget.history, hand it back here. When it is not empty, greeting is ignored.
showSourcesboolean
Show, under each answer, the knowledge cards it came from. Off by default.
tools('calc' | 'date' | ToolSpec)[]
Local tools whose results are handed to the model as facts: see the dedicated section.
theme / position / width / height / labelssee below
The widget’s look and wording: see Appearance & labels.

Appearance & labels

The widget follows your page instead of imposing its own look. Everything here is set per widget: two widgets on the same page can differ in theme, corner and colour.

embed({
  theme: 'auto',             // 'light' (default) | 'dark' | 'auto' — auto follows prefers-color-scheme, live
  position: 'bottom-left',   // 'bottom-right' (default) | 'bottom-left'
  width: 400, height: 600,   // px, clamped to 300-480 × 380-720; small screens still cap to the viewport
  labels: {                  // any language, any tone — keys you omit keep the lang defaults
    open: 'Chat öffnen',
    placeholder: 'Nachricht eingeben…',
    note: 'Lokale KI — läuft auf Ihrer GPU.',
    phases: { download: 'Modell wird geladen…' },
  },
});

lang covers English and French out of the box; labels is the door to every other language (open, close, placeholder, note, error, empty, help, sources, mb, phases). Labels are rendered as text, never as markup, and theme takes a keyword, not CSS: nothing an integrator passes here can inject styles or markup into the host page.

Controlling the widget

embed() returns a handle. It is what lets you unmount the widget — which matters in any app with client-side routing, where the widget would otherwise survive every route change and a second embed() would stack a second launcher on the page.

const widget = embed({ title: "Support" });

widget.open(); widget.close(); widget.toggle();
await widget.ask("Do you ship to Canada?");  // as if the visitor had typed it
widget.setKnowledge(newDocs);                // swap the cards, keep the conversation
widget.setHistory(saved);                    // resume a conversation
widget.history                               // Msg[]
widget.el                                    // the panel, for a style tweak
widget.destroy();                            // removes the DOM, cancels any generation in flight

destroy() leaves the engine loaded: the weights are shared by the page, so unmounting a widget never makes the next one download the model again. In React, the handle is exactly what an effect’s cleanup needs:

useEffect(() => {
  const widget = embed({ system: "You are our support agent." });
  return () => widget.destroy();
}, []);

Calling embed() on the server is harmless: you get an inert handle instead of a crash, so the same code can run on both sides.

Events

Both the widget handle and a session expose on(event, callback), which returns its own unsubscribe function. This is how you log conversations, measure engagement, and — most useful of all — learn that a visitor’s browser has no WebGPU, instead of that failure staying inside a chat bubble.

const off = widget.on('message', ({ role, content, sources }) => {
  analytics.track('chat', { role, content });
});
off();  // unsubscribe

widget.on('progress', (phase, p) => bar.value = p ? p.loaded / p.total : 0);
widget.on('ready',    () => console.log('model loaded'));
widget.on('open',     () => {});
widget.on('close',    () => {});
widget.on('error',    (err) => report(err));   // e.g. no WebGPU on this browser
widget.on('tool',     ({ name, result }) => {});  // a tool produced a result for this turn

phase is a stable key — init, download, tokenizer, gpu — never a sentence: you label it in your page’s own language. A listener that throws is caught and logged: your analytics can never break the widget. Sessions get ready, progress, message, error and tool (open and close are widget-only), and a session emits progress because ask() now preloads before its first turn: that first call used to download 149 MB with no way to say so.

The one failure a visitor can trigger without doing anything wrong carries a cause code: err.code === "no-webgpu" means this browser cannot run the assistant at all. Treat it as the signal to hide the widget rather than as a bug — status() answers the same question before anything is mounted. Errors on the message path are emitted AND thrown, so ask() keeps its own catch.

createSession(config?)

Headless: a conversation object for your own interface. Same config as embed() minus the visual options — including history to start from a stored conversation — plus temperature.

const session = createSession({ system: "You are a sommelier.", temperature: 0.7 });

const reply = await session.ask("A wine for oysters?", {
  onToken: (text) => output.textContent += text,  // streaming
  signal: controller.signal,                       // cancellable
});

session.history       // the Msg[] so far
session.lastSources   // the cards behind the last answer
session.setHistory(saved)      // resume a conversation
session.setKnowledge(newDocs)  // swap the cards, keep the conversation
session.on('message', log)     // same events as the widget
session.reset()  // same config, blank history
session.destroy()

setHistory() and setKnowledge() throw if a generation is running: finish or cancel the turn first. Both keep the engine and the weights untouched — swapping a catalogue does not cost a download, and no longer costs the conversation either. ask() also throws if a turn is already running on that session: one conversation, one turn at a time.

temperature defaults to 0.25 when you pass knowledge, and 0.55 otherwise — the same rule as the widget, and a measured one: at 0.55, reading one row out of a table went to the wrong column once in three. An assistant copying a figure out of a note has nothing to gain from sampling wide. Declaring temperature yourself still wins.

generate(options)

One shot: a prompt, a reply, no history kept. Takes the session config plus prompt, onToken, signal and onSources. It takes ONE object: called as generate("question", {…}) it throws a TypeError instead of quietly answering the string "undefined".

const answer = await generate({
  system: "Answer in one sentence.",
  prompt: "Why is the sky blue?",
  onToken: (text) => process(text),
});

Knowledge documents

Give the assistant your content: pages, FAQs, product sheets. Documents are chunked into passages in the browser, and only the 1–3 passages closest to the visitor’s question are given to the model. The ranking is local (lexical): nothing is sent anywhere.

knowledge: [
  "Plain strings work.",
  { title: "Shipping", text: "Free in France from 60 euros." },
],
knowledgeBudget: 800  // max tokens of passages per question

Tools

Give the assistant abilities beyond its notes: arithmetic, today’s date, or your own functions — a stock lookup, an order status, a cart total. The design is deliberate: the model NEVER decides to call a tool (below ~3B parameters, emitted tool calls are hallucinated — measured). Instead, detection is deterministic, your function runs in your page, and the model receives the result as a fact, exactly like a knowledge passage.

embed({
  tools: [
    'calc',   // detects arithmetic in the message and injects the exact result
    'date',   // the model knows today’s date (stable line in the system prompt)
    {
      name: 'stock',
      match: /stock|disponible/i,          // or a predicate: (question) => boolean
      run: async (question) => {           // runs in YOUR page — sync or async
        const n = await api.stockFor(question);
        return `${n} in stock`;
      },
    },
  ],
});

widget.on('tool', ({ name, result }) => trace(name, result));  // before generation, like onSources

The contract protects the visitor: a tool that throws, hangs (10 s cap) or returns nothing is simply absent from the turn — the turn itself never fails because of a tool. Results are capped at 600 characters: a result is a fact, not a report, and the default model has a short window. Nothing here touches the network unless YOUR run function does.

Sources of an answer

You can see which passages fed an answer. Two reasons this matters: a small model does get things wrong, and an answer a visitor can check is worth more than one merely asserted — and when yours answers oddly, this is how you tell a bad passage from a bad reading of a good one.

await session.ask(question, {
  onSources: (sources) => show(sources),  // before generation: the ranking is local and instant
});
session.lastSources  // [{ title, text, score, doc }]

embed({ knowledge: docs, showSources: true });          // in the widget, under each answer
widget.on('message', ({ sources }) => trace(sources));  // or without displaying anything

score is the lexical proximity to the question, doc the index of the document in your knowledge array, and the order is the order the passages were given to the model. An empty array is meaningful: no passage matched.

And what happens then depends on the message, because two situations must not be confused. A question asking for information gets the honest answer — “I do not have that information” — and that is the promise of the product. Anything else does not: “Are you ok?”, “PLEASE”, “hello” get a short, friendly reply. An assistant that stonewalls everything outside its notes is one visitors close, and it only takes a few such answers in a row for a small model to keep repeating them.

preload(), status() & runtime()

preload() downloads the engine and the model ahead of the first question: call it on a hover, or on the pricing page before support opens. onProgress receives the phase and, during download, the bytes: enough for a real progress bar.

await preload({
  onProgress: (status, p) => {
    if (p) bar.style.width = (100 * p.loaded / p.total) + "%";
  },
});

status()  // 'unavailable' (no WebGPU) | 'idle' | 'loading' | 'ready' | 'error'

status() answers synchronously: use it to decide whether to show the widget at all on browsers without WebGPU. runtime() reports where inference actually runs — "worker", "main", or "pending" before anything has started — which is what you check when you passed worker: true and want to know whether the fallback kicked in (a host CSP that forbids blob: makes it fall back to the main thread, silently and on purpose: a widget must never stop working over an execution choice).

One engine per page

The engine is a singleton per model URL: N widgets and N sessions on a page share one WebGPU init and one set of weights in VRAM. Mounting a second widget costs a DOM node, not 149 MB — and destroying one leaves the weights loaded for whoever comes next. This is also why worker and workerUrl only take effect before the first preload or ask: once the backend exists it is shared by the whole page, and a later embed() saying otherwise is ignored with a console warning rather than silently believed.

Versions & CDNs

Pin a version if you would rather the widget did not change under your feet:

https://brimkern.com/sdk-0.6.0.js   instead of   https://brimkern.com/sdk.js

<!-- or from the npm CDNs -->
https://unpkg.com/brimkern@0.6.0/dist/brimkern.iife.js
https://cdn.jsdelivr.net/npm/brimkern@0.6.0/dist/brimkern.iife.js

Servers, licence, links

Importing the package on a server does nothing until a browser runs it: Next.js, Remix and Astro pass without guards. MIT licence, like the whole engine.

npmjs.com/package/brimkern · GitHub · SDK page & live demo

Brimkern: open WebGPU engine, built by Romain Khanoyan. Local AI, WebGPU, on-device engines.