Xybrid
SDKs

Android

Native Android SDK for on-device ML inference

The Android SDK provides native Kotlin bindings to the Xybrid runtime via BoltFFI, enabling on-device ML inference in Android applications.

Installation

Add the Xybrid SDK from Maven Central:

build.gradle.kts
dependencies {
    implementation("ai.xybrid:xybrid-kotlin:0.10.1")
}

Minimum SDK: 24 (Android 7.0)

Quick Start

import ai.xybrid.*

class MyApplication : Application() {
    override fun onCreate() {
        super.onCreate()
        Xybrid.init(this)
    }
}

// Describe, then explicitly load a model from the registry
val model = Xybrid.model("kokoro-82m").load()

// Run inference
val result = model.runAsync(Envelope.text("Hello, world!"))

if (result.success) {
    println("Output: ${result.text}")
    println("Latency: ${result.latencyMs}ms")
}

Xybrid.init(context) also registers BatteryManager and (on API 29+) PowerManager.OnThermalStatusChangedListener observers, forwarding readings to the routing engine.

Initialization

Inference runs entirely on-device whether or not you authenticate. Pass an apiKey to light up the dashboard — that single call starts the telemetry exporter, and your model.run(...) calls automatically emit execution traces:

Xybrid.init(this, apiKey = BuildConfig.XYBRID_API_KEY)

Without a key, telemetry is disabled and the first inference logs a one-shot hint pointing at the dashboard (suppress with the XYBRID_QUIET=1 environment variable). Get a free key at dashboard.xybrid.dev. For a self-hosted dashboard, also pass ingestUrl.


Model Loading

From Registry

val model = Xybrid.model("whisper-tiny-ggml").load()

From Local Bundle

val model = Xybrid.model(ModelSource.bundle("/path/to/model.xyb")).load()

Xybrid.model(...) creates a cheap, unloaded reference. Use ModelSource.directory(path) and ModelSource.huggingFace("org/repo") for an unpacked directory and Hugging Face respectively. Suspending load() is the explicit boundary that may resolve metadata, download artifacts, access disk, and initialize the runtime. Use loadBlocking() only from Java or an existing worker thread.


Download Progress

load() blocks until the weights are on disk, so it has no progress to report while it runs. Start the download separately and watch it:

val loader = Xybrid.model("qwen3-0.6b")
loader.download()?.progress()?.collect { status ->
    bar.progress = (status.progress * 100).toInt()
    label.text = "${status.downloadedBytes} / ${status.totalBytes ?: 0}"
}
val model = loader.load()   // cached: returns at once

progress spans every file the model needs — weights plus companions such as a vision projector — and never moves backwards, so the bar neither restarts per file nor rewinds when a download retries. totalBytes is null when the source publishes no size (a Hugging Face repo); downloadedBytes is exact either way, which is what you need for megabytes, speed and time remaining.

Cancelling the consuming task only unsubscribes from updates. Call cancel() to stop the transfer itself — it takes effect within one chunk read and discards the partial file.


Input Envelopes

The Envelope factory creates type-safe inputs for different model types.

Audio (Speech Recognition)

val envelope = Envelope.audio(
    bytes = audioBytes,     // Raw PCM bytes
    sampleRate = 16000u,    // Sample rate in Hz
    channels = 1u           // Mono
)
val result = model.run(envelope)
println("Transcription: ${result.text}")

Text (Text-to-Speech)

// Simple text
val envelope = Envelope.text("Hello, how are you?")

// With voice and speed
val envelope = Envelope.text("Hello", voiceId = "af_heart", speed = 1.0)

val result = model.run(envelope)
val audioOutput = result.audioBytes  // Raw PCM audio

Embedding

val envelope = Envelope.embedding(listOf(0.1f, 0.2f, 0.3f))
val result = model.run(envelope)
val vector = result.embedding

Live Transcription

run() transcribes a finished audio buffer. A session transcribes speech as it arrives, which is what dictation and live captioning need.

Audio must be PCM float32, mono, 16 kHz — converting from the platform's microphone format is the caller's job.

val model = Xybrid.model("whisper-tiny").load()
val session = model.stream(streamingConfigWithVad(modelDir = vadModelDir, language = "en"))

scope.launch {
    session.partials().collect { partial -> textView.text = partial.text }
}

// From the AudioRecord loop, converted to float mono 16 kHz:
session.feed(pcm)

val transcript = session.flush()

Partial text is cumulative, not a delta: render each one in place of the previous rather than appending. A partial produced before you subscribe is delivered immediately, so the opening words of an utterance are never lost.

flush() ends the session and returns the full transcript — and reports an error if any chunk failed to transcribe, because that audio is gone and the transcript would otherwise look complete. cancel() ends the session and discards queued audio.

Voice-activity detection cuts chunks on speech boundaries instead of a fixed clock, which avoids the stutter you get when a window lands mid-word. It needs a Silero VAD model directory containing a model.onnx; none ships with the SDK, and without one the engine falls back to fixed windows.


Result Handling

val result = model.run(envelope)

if (result.success) {
    when (result.outputType) {
        "text" -> println("Text: ${result.text}")
        "audio" -> playAudio(result.audioBytes!!)
        "embedding" -> process(result.embedding!!)
    }
    println("Latency: ${result.latencyMs}ms")
    println("Latency: ${result.latencySeconds}s")
} else {
    println("Error: ${result.error}")
}

XybridResult Properties

PropertyTypeDescription
successBooleanWhether inference succeeded
errorString?Error message if failed
outputTypeString"text", "audio", or "embedding"
textString?Text output (ASR, LLM)
audioBytesByteArray?Audio output (TTS)
embeddingList<Float>?Embedding vector
latencyMsUIntInference latency in ms
isFailureBooleanConvenience: !success
latencySecondsDoubleLatency in seconds

Error Handling

The SDK uses sealed exception classes for type-safe error handling:

try {
    val model = Xybrid.model("kokoro-82m").load()
    val result = model.runAsync(envelope)
} catch (e: XybridException.ModelNotFound) {
    println("Model not found: ${e.modelId}")
} catch (e: XybridException.InferenceFailed) {
    println("Inference failed: ${e.message}")
} catch (e: XybridException.InvalidInput) {
    println("Invalid input: ${e.message}")
} catch (e: XybridException.IoException) {
    println("I/O error: ${e.message}")
} catch (e: XybridException) {
    // Catch-all with user-friendly message
    showError(e.displayMessage)
}

Platform Support

ArchitectureStatus
arm64-v8aSupported
armeabi-v7aSupported
x86_64Supported (emulator)

On this page