Unity
On-device ML inference SDK for Unity applications
The Unity SDK (ai.xybrid.sdk) provides C# bindings to the Xybrid runtime via C FFI, enabling on-device ML inference in Unity games and applications.
Installation
Install via Unity Package Manager using the git URL:
https://github.com/xybrid-ai/xybrid.git?path=bindings/unityIn Unity: Window > Package Manager > + > Add package from git URL
Minimum Unity version: 2021.3 LTS
Initialization
Initialize the SDK once at startup:
using Xybrid;
void Awake()
{
XybridClient.Initialize();
}Initialize() is thread-safe and idempotent — multiple calls are no-ops.
On Android, models are cached in Application.persistentDataPath/xybrid/models. To use another folder, call XybridBolt.XybridBolt.InitSdkCacheDir(path) before loading or listing any model; the first folder set wins.
Inference runs entirely on-device whether or not you authenticate. Pass an apiKey to light up the dashboard — that single call starts the telemetry exporter, and your model.Run(...) calls automatically emit execution traces:
XybridClient.Initialize(apiKey: "xy_live_...");Without a key, telemetry is disabled and the first inference logs a one-shot hint pointing at the dashboard (suppress with the XYBRID_QUIET=1 environment variable). Get a free key at dashboard.xybrid.dev. For a self-hosted dashboard, also pass ingestUrl. For advanced telemetry settings (batch size, flush interval, device attributes), build a TelemetryConfig and call XybridClient.InitializeTelemetry instead.
Quick Start
using Xybrid;
// Load and run in one line
var model = XybridClient.LoadModel("kokoro-82m");
using (var envelope = Envelope.Text("Halt! None shall pass without the king's seal."))
using (var result = model.Run(envelope))
{
result.ThrowIfFailed();
Debug.Log($"Output: {result.Text}");
Debug.Log($"Latency: {result.LatencyMs}ms");
}Model Loading
From Registry
using (var loader = ModelLoader.FromRegistry("whisper-tiny"))
{
var model = loader.Load();
// Use model for inference...
}From Local Bundle
using (var loader = ModelLoader.FromBundle("path/to/model.xyb"))
{
var model = loader.Load();
}Convenience Methods
// Load directly without managing the loader
var model = XybridClient.LoadModel("kokoro-82m");
var model = XybridClient.LoadModelFromBundle("path/to/model.xyb");Download Progress
load() blocks until the weights are on disk, so it has no progress to report
while it runs. Start the download separately and watch it:
using var download = ModelLoader.FromRegistry("qwen3-0.6b").StartDownload();
while (!download.IsFinished())
{
// Status() never blocks, so it is safe once per frame.
slider.value = download.Status().Progress;
yield return null;
}
using var loader = ModelLoader.FromRegistry("qwen3-0.6b");
var model = loader.Load(); // cached: returns at onceprogress spans every file the model needs — weights plus companions such as a
vision projector — and never moves backwards, so the bar neither restarts per
file nor rewinds when a download retries. totalBytes is null when the source
publishes no size (a Hugging Face repo); downloadedBytes is exact either way,
which is what you need for megabytes, speed and time remaining.
Cancelling the consuming task only unsubscribes from updates. Call cancel()
to stop the transfer itself — it takes effect within one chunk read and
discards the partial file.
Input Envelopes
Text
using (var envelope = Envelope.Text("The dragon sleeps in the northern cave."))
{
using (var result = model.Run(envelope))
{
Debug.Log(result.Text);
}
}Text with Role (for conversations)
using (var envelope = Envelope.Text("You are a blacksmith in a medieval village.", MessageRole.System))
{
// Use with ConversationContext
}Audio
byte[] audioBytes = File.ReadAllBytes("audio.wav");
using (var envelope = Envelope.Audio(audioBytes, sampleRate: 16000, channels: 1))
{
using (var result = model.Run(envelope))
{
Debug.Log($"Transcription: {result.Text}");
}
}Convenience Methods
// Text-in, text-out shorthand
string output = model.RunText("Tell me about the ancient ruins.");
// Audio-in, text-out shorthand (player voice command)
string transcription = model.RunAudio(microphoneBytes, sampleRate: 16000);Live Transcription
run() transcribes a finished audio buffer. A session transcribes speech as
it arrives, which is what dictation and live captioning need.
Audio must be PCM float32, mono, 16 kHz — converting from the platform's microphone format is the caller's job.
using var session = model.Stream(StreamingConfigs.VoiceActivity(vadModelDir, language: "en"));
// Drive the UI from the partial stream.
await foreach (var partial in session.Partials(cancellationToken))
{
label.text = partial.Text;
}
// From the Microphone clip, converted to float mono 16 kHz:
session.Feed(pcm);
string transcript = session.Flush();Partial text is cumulative, not a delta: render each one in place of the previous rather than appending. A partial produced before you subscribe is delivered immediately, so the opening words of an utterance are never lost.
flush() ends the session and returns the full transcript — and reports an
error if any chunk failed to transcribe, because that audio is gone and the
transcript would otherwise look complete. cancel() ends the session and
discards queued audio.
Voice-activity detection cuts chunks on speech boundaries instead of a fixed
clock, which avoids the stutter you get when a window lands mid-word. It needs
a Silero VAD model directory containing a model.onnx; none ships with the
SDK, and without one the engine falls back to fixed windows.
Result Handling
using (var result = model.Run(envelope))
{
if (result.Success)
{
Debug.Log($"Output: {result.Text}");
Debug.Log($"Latency: {result.LatencyMs}ms");
}
else
{
Debug.LogError($"Error: {result.Error}");
}
}
// Or throw on failure
using (var result = model.Run(envelope))
{
result.ThrowIfFailed();
// Safe to use result.Text here
}InferenceResult Properties
| Property | Type | Description |
|---|---|---|
Success | bool | Whether inference succeeded |
Error | string | Error message if failed |
Text | string | Text output (ASR, LLM) |
LatencyMs | uint | Inference latency in ms |
Conversation Context
Multi-turn LLM conversations with automatic history management:
using (var context = new ConversationContext())
{
context.SetSystem("You are a wise elder in a forest village. You speak in short, cryptic phrases.");
context.SetMaxHistoryLength(50); // FIFO pruning after 50 messages
// First turn
context.Push("Where is the lost temple?", MessageRole.User);
using (var envelope = Envelope.Text("Where is the lost temple?"))
using (var result = model.Run(envelope, context))
{
result.ThrowIfFailed();
context.Push(result.Text, MessageRole.Assistant);
Debug.Log(result.Text);
}
// Second turn (has full history)
context.Push("What dangers await there?", MessageRole.User);
using (var envelope = Envelope.Text("What dangers await there?"))
using (var result = model.Run(envelope, context))
{
result.ThrowIfFailed();
context.Push(result.Text, MessageRole.Assistant);
Debug.Log(result.Text);
}
// Clear history but keep system prompt
context.Clear();
}ConversationContext Properties
| Property | Type | Description |
|---|---|---|
Id | string | Conversation UUID |
HistoryLength | uint | Current message count |
HasSystem | bool | System prompt set? |
LLM Streaming
Unity supports token-by-token streaming for LLM models via callbacks:
// Check if the model supports streaming
if (model.SupportsTokenStreaming)
{
using (var envelope = Envelope.Text("Tell me a story about a brave knight."))
using (var result = model.RunStreaming(envelope, token =>
{
// Called for each token as it's generated
Debug.Log(token.Token);
}))
{
result.ThrowIfFailed();
Debug.Log($"Full response: {result.Text}");
}
}Streaming with Conversation Context
using (var context = new ConversationContext())
{
context.SetSystem("You are a tavern keeper in a fantasy RPG.");
context.Push("What's on the menu today?", MessageRole.User);
using (var envelope = Envelope.Text("What's on the menu today?"))
using (var result = model.RunStreaming(envelope, context, token =>
{
Debug.Log(token.Token); // Each token as it arrives
}))
{
result.ThrowIfFailed();
context.Push(result.Text, MessageRole.Assistant);
}
}Convenience Streaming
// Text-in, streaming text-out
string response = model.RunStreamingText("Describe the dungeon entrance.", token =>
{
// Update UI with each token
dialogueText.text += token.Token;
});IDisposable Pattern
All resource-holding classes implement IDisposable. Use using blocks to ensure cleanup:
using (var loader = ModelLoader.FromRegistry("kokoro-82m"))
using (var model = loader.Load())
using (var envelope = Envelope.Text("Your quest is complete, hero."))
using (var result = model.Run(envelope))
{
result.ThrowIfFailed();
Debug.Log(result.Text);
}Classes with IDisposable: Model, ModelLoader, Envelope, InferenceResult, ConversationContext
Error Handling
try
{
var model = XybridClient.LoadModel("nonexistent-model");
}
catch (ModelNotFoundException e)
{
Debug.LogError($"Model not found: {e.ModelId}");
}
catch (InferenceException e)
{
Debug.LogError($"Inference failed: {e.Message}");
}
catch (XybridException e)
{
Debug.LogError($"SDK error: {e.Message}");
}Exception Hierarchy
| Exception | Description |
|---|---|
XybridException | Base exception for all SDK errors |
ModelNotFoundException | Model ID not found in registry (has ModelId property) |
InferenceException | Inference execution failed |
Platform Support
| Platform | Status |
|---|---|
| macOS (Apple Silicon) | Supported |
| macOS (Intel) | Supported |
| Windows | Planned |
| Linux | Planned |