Xybrid
SDKs

Web

Browser SDK preview for on-device ML inference with WebAssembly and WebGPU

The Web SDK (@xybrid/web) runs GGUF language models locally in a browser through Xybrid's existing Rust llama wrappers and llama.cpp, compiled to WebAssembly. A worker streams tokens, supports cancellation, and reports memory and speed. CPU execution uses WASM SIMD; WebGPU requires shader-f16.

The package is a private ESM preview in bindings/web, not yet published to npm. It replaces the former LiteRT tensor/text package. The complete native SDK API and platform services are still being integrated.

Build and try it

Install Bun, pnpm, and Bazelisk, then:

git submodule update --init vendor/llama-cpp
cd bindings/web
pnpm install --frozen-lockfile
pnpm build
pnpm dev:example

Bazel downloads pinned Rust and Emscripten toolchains and builds CPU and WebGPU variants. The example verifies and loads a 91.7 MB SmolLM2-135M-Instruct Q4_0 GGUF. For a static site, run pnpm build:example and serve example/dist over HTTPS or localhost.

Load and stream

Copy dist/runtime/ into the application's public /xybrid/runtime/ directory. Keep the generated worker with the package. Bundlers import @xybrid/web; direct browser imports can serve the complete dist/ directory under /sdk/, import /sdk/index.js, and use wasmPath: "/sdk/runtime".

import { XybridLlm } from "@xybrid/web";

const model = await XybridLlm.fromUrl("/models/model.gguf", {
  wasmPath: "/xybrid/runtime",
  accelerator: "auto",
  contextLength: 512,
});
try {
  for await (const delta of model.generateStream("Tell me a short story.", {
    maxOutputTokens: 64,
  })) {
    output.append(delta);
  }
  console.log(model.loaded, model.lastRun);
} finally {
  await model.dispose();
  console.log(model.releasedMemory);
}

generate() returns a complete string. Breaking the iterator cancels decoding and waits for native work to settle. cancel() provides an explicit stop; dispose() cancels, releases Rust allocations, and terminates the worker. Overlapping generations on one model are rejected.

Model sources

FactoryInput
fromUrl(url, options?)Direct GGUF URL, optionally pinned with sizeBytes and sha256
load(url, options?)Xybrid metadata with a Gguf execution template
fromRegistry(id, options?)Registry resolution in gguf format
fromHuggingFace(repo, options?)A top-level GGUF selected with file and optional revision
const model = await XybridLlm.fromHuggingFace(
  "QuantFactory/SmolLM2-135M-Instruct-GGUF",
  { file: "SmolLM2-135M-Instruct.Q4_0.gguf", accelerator: "wasm" },
);

Registry models require an HTTPS download URL, declared size, SHA-256, and compatible metadata. Hugging Face models enforce size and verify a valid Git-LFS SHA-256 OID when present. A direct load can supply the same integrity pin. Load options accept signal and onDownloadProgress.

Metadata uses execution_template.type: "Gguf", a listed bare .gguf filename, and optional context_length. Model size is capped at 512 MiB, metadata at 1 MiB, and contexts at 32,768 tokens. Runtime assets must be on the page's origin; model servers need CORS. Only the selected CPU or GPU runtime downloads.

Measurements and current limits

loaded records Rust version, backend, model bytes, download/initialization/load timings, and WASM memory. lastRun records token counts, first-token latency, decode throughput, cancellation, and sampled memory peaks. releasedMemory reports allocations after Rust destruction. These measurements exclude JavaScript buffers, browser process memory, GPU allocations, and transient peaks inside model loading or prefill.

Generation uses an embedded GGUF chat template, greedy decoding, a fresh user turn, and a cleared KV cache. Direct loads default to a 512-token context. Cancellation occurs at token boundaries; prefill runs synchronously in the worker. External chat templates, metadata sampling/processing, tensor inference, vision, speech, and pipelines are not supported by this preview.

CI builds both backends and tests the packaged production SDK in Chromium. Build sizes and benchmarks are uploaded as artifacts. Actual GPU inference is an opt-in test on compatible hardware, or the manually dispatched WebGPU browser inference workflow on a self-hosted runner labelled webgpu.

For asset layout, migration details, and verification commands, see the package README.

On this page