Web
Browser SDK preview for on-device ML inference with WebAssembly and WebGPU
The Web SDK (@xybrid/web) runs GGUF language models locally in a browser
through Xybrid's existing Rust llama wrappers and llama.cpp, compiled to
WebAssembly. A worker streams tokens, supports cancellation, and reports memory
and speed. CPU execution uses WASM SIMD; WebGPU requires shader-f16.
The package is a private ESM preview in
bindings/web,
not yet published to npm. It replaces the former LiteRT tensor/text package.
The complete native SDK API and platform services are still being integrated.
Build and try it
Install Bun, pnpm, and Bazelisk, then:
git submodule update --init vendor/llama-cpp
cd bindings/web
pnpm install --frozen-lockfile
pnpm build
pnpm dev:exampleBazel downloads pinned Rust and Emscripten toolchains and builds CPU and WebGPU
variants. The example verifies and loads a 91.7 MB SmolLM2-135M-Instruct Q4_0
GGUF. For a static site, run pnpm build:example and serve example/dist over
HTTPS or localhost.
Load and stream
Copy dist/runtime/ into the application's public /xybrid/runtime/ directory.
Keep the generated worker with the package. Bundlers import @xybrid/web;
direct browser imports can serve the complete dist/ directory under /sdk/,
import /sdk/index.js, and use wasmPath: "/sdk/runtime".
import { XybridLlm } from "@xybrid/web";
const model = await XybridLlm.fromUrl("/models/model.gguf", {
wasmPath: "/xybrid/runtime",
accelerator: "auto",
contextLength: 512,
});
try {
for await (const delta of model.generateStream("Tell me a short story.", {
maxOutputTokens: 64,
})) {
output.append(delta);
}
console.log(model.loaded, model.lastRun);
} finally {
await model.dispose();
console.log(model.releasedMemory);
}generate() returns a complete string. Breaking the iterator cancels decoding
and waits for native work to settle. cancel() provides an explicit stop;
dispose() cancels, releases Rust allocations, and terminates the worker.
Overlapping generations on one model are rejected.
Model sources
| Factory | Input |
|---|---|
fromUrl(url, options?) | Direct GGUF URL, optionally pinned with sizeBytes and sha256 |
load(url, options?) | Xybrid metadata with a Gguf execution template |
fromRegistry(id, options?) | Registry resolution in gguf format |
fromHuggingFace(repo, options?) | A top-level GGUF selected with file and optional revision |
const model = await XybridLlm.fromHuggingFace(
"QuantFactory/SmolLM2-135M-Instruct-GGUF",
{ file: "SmolLM2-135M-Instruct.Q4_0.gguf", accelerator: "wasm" },
);Registry models require an HTTPS download URL, declared size, SHA-256, and
compatible metadata. Hugging Face models enforce size and verify a valid Git-LFS
SHA-256 OID when present. A direct load can supply the same integrity pin. Load
options accept signal and onDownloadProgress.
Metadata uses execution_template.type: "Gguf", a listed bare .gguf filename,
and optional context_length. Model size is capped at 512 MiB, metadata at 1 MiB,
and contexts at 32,768 tokens. Runtime assets must be on the page's origin;
model servers need CORS. Only the selected CPU or GPU runtime downloads.
Measurements and current limits
loaded records Rust version, backend, model bytes, download/initialization/load
timings, and WASM memory. lastRun records token counts, first-token latency,
decode throughput, cancellation, and sampled memory peaks. releasedMemory
reports allocations after Rust destruction. These measurements exclude
JavaScript buffers, browser process memory, GPU allocations, and transient peaks
inside model loading or prefill.
Generation uses an embedded GGUF chat template, greedy decoding, a fresh user turn, and a cleared KV cache. Direct loads default to a 512-token context. Cancellation occurs at token boundaries; prefill runs synchronously in the worker. External chat templates, metadata sampling/processing, tensor inference, vision, speech, and pipelines are not supported by this preview.
CI builds both backends and tests the packaged production SDK in Chromium.
Build sizes and benchmarks are uploaded as artifacts. Actual GPU inference is
an opt-in test on compatible hardware, or the manually dispatched WebGPU browser inference workflow on a self-hosted runner labelled webgpu.
For asset layout, migration details, and verification commands, see the package README.