Xybrid
SDKs

Unity

On-device ML inference SDK for Unity applications

The Unity SDK (ai.xybrid.sdk) provides C# bindings to the Xybrid runtime via C FFI, enabling on-device ML inference in Unity games and applications.

Installation

Install via Unity Package Manager using the git URL:

https://github.com/xybrid-ai/xybrid.git?path=bindings/unity

In Unity: Window > Package Manager > + > Add package from git URL

Minimum Unity version: 2021.3 LTS

Initialization

Initialize the SDK once at startup:

using Xybrid;

void Awake()
{
    XybridClient.Initialize();
}

Initialize() is thread-safe and idempotent — multiple calls are no-ops.

On Android, models are cached in Application.persistentDataPath/xybrid/models. To use another folder, call XybridBolt.XybridBolt.InitSdkCacheDir(path) before loading or listing any model; the first folder set wins.

Inference runs entirely on-device whether or not you authenticate. Pass an apiKey to light up the dashboard — that single call starts the telemetry exporter, and your model.Run(...) calls automatically emit execution traces:

XybridClient.Initialize(apiKey: "xy_live_...");

Without a key, telemetry is disabled and the first inference logs a one-shot hint pointing at the dashboard (suppress with the XYBRID_QUIET=1 environment variable). Get a free key at dashboard.xybrid.dev. For a self-hosted dashboard, also pass ingestUrl. For advanced telemetry settings (batch size, flush interval, device attributes), build a TelemetryConfig and call XybridClient.InitializeTelemetry instead.

Quick Start

using Xybrid;

// Load and run in one line
var model = XybridClient.LoadModel("kokoro-82m");
using (var envelope = Envelope.Text("Halt! None shall pass without the king's seal."))
using (var result = model.Run(envelope))
{
    result.ThrowIfFailed();
    Debug.Log($"Output: {result.Text}");
    Debug.Log($"Latency: {result.LatencyMs}ms");
}

Model Loading

From Registry

using (var loader = ModelLoader.FromRegistry("whisper-tiny"))
{
    var model = loader.Load();
    // Use model for inference...
}

From Local Bundle

using (var loader = ModelLoader.FromBundle("path/to/model.xyb"))
{
    var model = loader.Load();
}

Convenience Methods

// Load directly without managing the loader
var model = XybridClient.LoadModel("kokoro-82m");
var model = XybridClient.LoadModelFromBundle("path/to/model.xyb");

Download Progress

load() blocks until the weights are on disk, so it has no progress to report while it runs. Start the download separately and watch it:

using var download = ModelLoader.FromRegistry("qwen3-0.6b").StartDownload();
while (!download.IsFinished())
{
    // Status() never blocks, so it is safe once per frame.
    slider.value = download.Status().Progress;
    yield return null;
}

using var loader = ModelLoader.FromRegistry("qwen3-0.6b");
var model = loader.Load();   // cached: returns at once

progress spans every file the model needs — weights plus companions such as a vision projector — and never moves backwards, so the bar neither restarts per file nor rewinds when a download retries. totalBytes is null when the source publishes no size (a Hugging Face repo); downloadedBytes is exact either way, which is what you need for megabytes, speed and time remaining.

Cancelling the consuming task only unsubscribes from updates. Call cancel() to stop the transfer itself — it takes effect within one chunk read and discards the partial file.


Input Envelopes

Text

using (var envelope = Envelope.Text("The dragon sleeps in the northern cave."))
{
    using (var result = model.Run(envelope))
    {
        Debug.Log(result.Text);
    }
}

Text with Role (for conversations)

using (var envelope = Envelope.Text("You are a blacksmith in a medieval village.", MessageRole.System))
{
    // Use with ConversationContext
}

Audio

byte[] audioBytes = File.ReadAllBytes("audio.wav");
using (var envelope = Envelope.Audio(audioBytes, sampleRate: 16000, channels: 1))
{
    using (var result = model.Run(envelope))
    {
        Debug.Log($"Transcription: {result.Text}");
    }
}

Convenience Methods

// Text-in, text-out shorthand
string output = model.RunText("Tell me about the ancient ruins.");

// Audio-in, text-out shorthand (player voice command)
string transcription = model.RunAudio(microphoneBytes, sampleRate: 16000);

Live Transcription

run() transcribes a finished audio buffer. A session transcribes speech as it arrives, which is what dictation and live captioning need.

Audio must be PCM float32, mono, 16 kHz — converting from the platform's microphone format is the caller's job.

using var session = model.Stream(StreamingConfigs.VoiceActivity(vadModelDir, language: "en"));

// Drive the UI from the partial stream.
await foreach (var partial in session.Partials(cancellationToken))
{
    label.text = partial.Text;
}

// From the Microphone clip, converted to float mono 16 kHz:
session.Feed(pcm);

string transcript = session.Flush();

Partial text is cumulative, not a delta: render each one in place of the previous rather than appending. A partial produced before you subscribe is delivered immediately, so the opening words of an utterance are never lost.

flush() ends the session and returns the full transcript — and reports an error if any chunk failed to transcribe, because that audio is gone and the transcript would otherwise look complete. cancel() ends the session and discards queued audio.

Voice-activity detection cuts chunks on speech boundaries instead of a fixed clock, which avoids the stutter you get when a window lands mid-word. It needs a Silero VAD model directory containing a model.onnx; none ships with the SDK, and without one the engine falls back to fixed windows.


Result Handling

using (var result = model.Run(envelope))
{
    if (result.Success)
    {
        Debug.Log($"Output: {result.Text}");
        Debug.Log($"Latency: {result.LatencyMs}ms");
    }
    else
    {
        Debug.LogError($"Error: {result.Error}");
    }
}

// Or throw on failure
using (var result = model.Run(envelope))
{
    result.ThrowIfFailed();
    // Safe to use result.Text here
}

InferenceResult Properties

PropertyTypeDescription
SuccessboolWhether inference succeeded
ErrorstringError message if failed
TextstringText output (ASR, LLM)
LatencyMsuintInference latency in ms

Conversation Context

Multi-turn LLM conversations with automatic history management:

using (var context = new ConversationContext())
{
    context.SetSystem("You are a wise elder in a forest village. You speak in short, cryptic phrases.");
    context.SetMaxHistoryLength(50); // FIFO pruning after 50 messages

    // First turn
    context.Push("Where is the lost temple?", MessageRole.User);
    using (var envelope = Envelope.Text("Where is the lost temple?"))
    using (var result = model.Run(envelope, context))
    {
        result.ThrowIfFailed();
        context.Push(result.Text, MessageRole.Assistant);
        Debug.Log(result.Text);
    }

    // Second turn (has full history)
    context.Push("What dangers await there?", MessageRole.User);
    using (var envelope = Envelope.Text("What dangers await there?"))
    using (var result = model.Run(envelope, context))
    {
        result.ThrowIfFailed();
        context.Push(result.Text, MessageRole.Assistant);
        Debug.Log(result.Text);
    }

    // Clear history but keep system prompt
    context.Clear();
}

ConversationContext Properties

PropertyTypeDescription
IdstringConversation UUID
HistoryLengthuintCurrent message count
HasSystemboolSystem prompt set?

LLM Streaming

Unity supports token-by-token streaming for LLM models via callbacks:

// Check if the model supports streaming
if (model.SupportsTokenStreaming)
{
    using (var envelope = Envelope.Text("Tell me a story about a brave knight."))
    using (var result = model.RunStreaming(envelope, token =>
    {
        // Called for each token as it's generated
        Debug.Log(token.Token);
    }))
    {
        result.ThrowIfFailed();
        Debug.Log($"Full response: {result.Text}");
    }
}

Streaming with Conversation Context

using (var context = new ConversationContext())
{
    context.SetSystem("You are a tavern keeper in a fantasy RPG.");

    context.Push("What's on the menu today?", MessageRole.User);
    using (var envelope = Envelope.Text("What's on the menu today?"))
    using (var result = model.RunStreaming(envelope, context, token =>
    {
        Debug.Log(token.Token);  // Each token as it arrives
    }))
    {
        result.ThrowIfFailed();
        context.Push(result.Text, MessageRole.Assistant);
    }
}

Convenience Streaming

// Text-in, streaming text-out
string response = model.RunStreamingText("Describe the dungeon entrance.", token =>
{
    // Update UI with each token
    dialogueText.text += token.Token;
});

IDisposable Pattern

All resource-holding classes implement IDisposable. Use using blocks to ensure cleanup:

using (var loader = ModelLoader.FromRegistry("kokoro-82m"))
using (var model = loader.Load())
using (var envelope = Envelope.Text("Your quest is complete, hero."))
using (var result = model.Run(envelope))
{
    result.ThrowIfFailed();
    Debug.Log(result.Text);
}

Classes with IDisposable: Model, ModelLoader, Envelope, InferenceResult, ConversationContext


Error Handling

try
{
    var model = XybridClient.LoadModel("nonexistent-model");
}
catch (ModelNotFoundException e)
{
    Debug.LogError($"Model not found: {e.ModelId}");
}
catch (InferenceException e)
{
    Debug.LogError($"Inference failed: {e.Message}");
}
catch (XybridException e)
{
    Debug.LogError($"SDK error: {e.Message}");
}

Exception Hierarchy

ExceptionDescription
XybridExceptionBase exception for all SDK errors
ModelNotFoundExceptionModel ID not found in registry (has ModelId property)
InferenceExceptionInference execution failed

Platform Support

PlatformStatus
macOS (Apple Silicon)Supported
macOS (Intel)Supported
WindowsPlanned
LinuxPlanned

On this page