Xybrid
SDKs

iOS / macOS

Native Swift SDK for on-device ML inference

The Swift SDK provides native bindings to the Xybrid runtime via BoltFFI for iOS and macOS applications.

The Swift SDK is in early access. Releases ship a prebuilt XybridFFI xcframework via Swift Package Manager. Known issue: building for the iOS Simulator on Apple Silicon requires the useLocalNatives = true workaround — see #179. For cross-platform development, see the Flutter SDK.

Installation

Add the Xybrid package via Swift Package Manager:

Package.swift
dependencies: [
    .package(url: "https://github.com/xybrid-ai/xybrid", from: "0.10.1")
]

Or in Xcode: File > Add Package Dependencies and enter the repository URL.

Supported platforms: iOS 13.0+, macOS 10.15+

Quick Start

import Xybrid

@main
struct MyApp: App {
    init() { Xybrid.initialize() }
    var body: some Scene { /* ... */ }
}

// Describe, then explicitly load a model from the registry
let model = try await Xybrid.model("kokoro-82m").load()

// Run inference
let envelope = Envelope.text("Hello, world!")
let result = try await model.runAsync(envelope: envelope)

if result.success {
    print("Output: \(result.text ?? "")")
    print("Latency: \(result.latency)s")
}

On iOS, Xybrid.initialize() enables UIDevice battery monitoring and subscribes to UIDevice.batteryLevelDidChangeNotification, forwarding readings to the routing engine. Thermal state on Apple platforms is sourced from NSProcessInfo.thermalState directly in xybrid-core.

Initialization

Inference runs entirely on-device whether or not you authenticate. Pass an apiKey to light up the dashboard — that single call starts the telemetry exporter, and your model.run(...) calls automatically emit execution traces:

Xybrid.initialize(
    apiKey: ProcessInfo.processInfo.environment["XYBRID_API_KEY"]
)

Without a key, telemetry is disabled and the first inference logs a one-shot hint pointing at the dashboard (suppress with the XYBRID_QUIET=1 environment variable). Get a free key at dashboard.xybrid.dev. For a self-hosted dashboard, also pass ingestUrl.


Model Loading

From Registry

let model = try await Xybrid.model("whisper-tiny-ggml").load()

From Local Bundle

let model = try await Xybrid.model(.bundle(bundleURL)).load()

Xybrid.model(...) creates a cheap, unloaded reference. Local bundle and directory sources use URL through .bundle(url) and .directory(url); Hugging Face uses .huggingFace("org/repo"). Async load() is the explicit boundary for resolving, downloading, disk access, and runtime initialization. Use loadSync() only from an existing worker thread.


Download Progress

load() blocks until the weights are on disk, so it has no progress to report while it runs. Start the download separately and watch it:

let loader = Xybrid.model("qwen3-0.6b")
if let download = loader.download() {
    for await status in download.progress() {
        bar.progress = Float(status.progress)
        label.text = "\(status.downloadedBytes) / \(status.totalBytes ?? 0)"
    }
}
let model = try await loader.load()   // cached: returns at once

progress spans every file the model needs — weights plus companions such as a vision projector — and never moves backwards, so the bar neither restarts per file nor rewinds when a download retries. totalBytes is null when the source publishes no size (a Hugging Face repo); downloadedBytes is exact either way, which is what you need for megabytes, speed and time remaining.

Cancelling the consuming task only unsubscribes from updates. Call cancel() to stop the transfer itself — it takes effect within one chunk read and discards the partial file.


Input Envelopes

Audio (Speech Recognition)

let envelope = Envelope.audio(pcmData: audioData, sampleRate: 16000, channels: 1)
let result = try model.run(envelope: envelope)
print("Transcription: \(result.text ?? "")")

Text (Text-to-Speech)

// Simple text
let envelope = Envelope.text("Hello, how are you?")

// With voice and speed
let envelope = Envelope.text("Hello", voice: "af_heart", speed: 1.0)

let result = try model.run(envelope: envelope)
if let audioBytes = result.audioBytes {
    // Play or save audio
}

Embedding

let envelope = Envelope.embedding(data: [0.1, 0.2, 0.3])
let result = try model.run(envelope: envelope)
if let vector = result.embedding {
    // Process embedding vector
}

Live Transcription

run() transcribes a finished audio buffer. A session transcribes speech as it arrives, which is what dictation and live captioning need.

Audio must be PCM float32, mono, 16 kHz — converting from the platform's microphone format is the caller's job.

let model = try await Xybrid.model("whisper-tiny").load()
let session = try model.stream(config: .voiceActivity(modelDir: vadModelDir, language: "en"))

Task {
    for await partial in session.partials() {
        label.text = partial.text
    }
}

// From the AVAudioEngine tap, converted to 16 kHz mono Float32:
try session.feed(samples: pcm)

// When the user stops talking:
let transcript = try session.flush()

Partial text is cumulative, not a delta: render each one in place of the previous rather than appending. A partial produced before you subscribe is delivered immediately, so the opening words of an utterance are never lost.

flush() ends the session and returns the full transcript — and reports an error if any chunk failed to transcribe, because that audio is gone and the transcript would otherwise look complete. cancel() ends the session and discards queued audio.

Voice-activity detection cuts chunks on speech boundaries instead of a fixed clock, which avoids the stutter you get when a window lands mid-word. It needs a Silero VAD model directory containing a model.onnx; none ships with the SDK, and without one the engine falls back to fixed windows.


Result Handling

let result = try model.run(envelope: envelope)

if result.success {
    switch result.outputType {
    case "text":
        print("Text: \(result.text!)")
    case "audio":
        playAudio(result.audioBytes!)
    case "embedding":
        process(result.embedding!)
    default:
        break
    }
    print("Latency: \(result.latency)s")  // TimeInterval in seconds
} else {
    print("Error: \(result.error ?? "Unknown")")
}

XybridResult Properties

PropertyTypeDescription
successBoolWhether inference succeeded
errorString?Error message if failed
outputTypeString"text", "audio", or "embedding"
textString?Text output (ASR, LLM)
audioBytesData?Audio output (TTS)
embedding[Float]?Embedding vector
latencyMsUInt32Inference latency in ms
isFailureBoolConvenience: !success
latencyTimeIntervalLatency in seconds

Error Handling

The SDK uses a Swift error enum for type-safe error handling:

do {
    let model = try await Xybrid.model("kokoro-82m").load()
    let result = try await model.runAsync(envelope: envelope)
} catch XybridError.ModelNotFound(let modelId) {
    print("Model not found: \(modelId)")
} catch XybridError.InferenceFailed(let message) {
    print("Inference failed: \(message)")
} catch XybridError.InvalidInput(let message) {
    print("Invalid input: \(message)")
} catch XybridError.IoError(let message) {
    print("I/O error: \(message)")
} catch {
    print("Unexpected error: \(error.localizedDescription)")
}

Type Aliases

The SDK provides short aliases for convenience:

AliasFull Type
ModelXybridModel
EnvelopeXybridEnvelope
ResultXybridResult
VoiceInfoXybridVoiceInfo
GenerationConfigXybridGenerationConfig
XybridSDKErrorXybridError

ModelLoader is a cheap model reference. Construct it through Xybrid.model(...), then call load() explicitly before inference.


Platform Support

PlatformStatusAccelerators
iOSSupportedMetal, CoreML, ANE
macOS (Apple Silicon)SupportedMetal, CoreML, ANE
macOS (Intel)SupportedCPU

On this page