iOS / macOS
Native Swift SDK for on-device ML inference
The Swift SDK provides native bindings to the Xybrid runtime via BoltFFI for iOS and macOS applications.
The Swift SDK is in early access. Releases ship a prebuilt XybridFFI xcframework via Swift Package Manager. Known issue: building for the iOS Simulator on Apple Silicon requires the useLocalNatives = true workaround — see #179. For cross-platform development, see the Flutter SDK.
Installation
Add the Xybrid package via Swift Package Manager:
dependencies: [
.package(url: "https://github.com/xybrid-ai/xybrid", from: "0.10.1")
]Or in Xcode: File > Add Package Dependencies and enter the repository URL.
Supported platforms: iOS 13.0+, macOS 10.15+
Quick Start
import Xybrid
@main
struct MyApp: App {
init() { Xybrid.initialize() }
var body: some Scene { /* ... */ }
}
// Describe, then explicitly load a model from the registry
let model = try await Xybrid.model("kokoro-82m").load()
// Run inference
let envelope = Envelope.text("Hello, world!")
let result = try await model.runAsync(envelope: envelope)
if result.success {
print("Output: \(result.text ?? "")")
print("Latency: \(result.latency)s")
}On iOS, Xybrid.initialize() enables UIDevice battery monitoring and subscribes to UIDevice.batteryLevelDidChangeNotification, forwarding readings to the routing engine. Thermal state on Apple platforms is sourced from NSProcessInfo.thermalState directly in xybrid-core.
Initialization
Inference runs entirely on-device whether or not you authenticate. Pass an apiKey to light up the dashboard — that single call starts the telemetry exporter, and your model.run(...) calls automatically emit execution traces:
Xybrid.initialize(
apiKey: ProcessInfo.processInfo.environment["XYBRID_API_KEY"]
)Without a key, telemetry is disabled and the first inference logs a one-shot hint pointing at the dashboard (suppress with the XYBRID_QUIET=1 environment variable). Get a free key at dashboard.xybrid.dev. For a self-hosted dashboard, also pass ingestUrl.
Model Loading
From Registry
let model = try await Xybrid.model("whisper-tiny-ggml").load()From Local Bundle
let model = try await Xybrid.model(.bundle(bundleURL)).load()Xybrid.model(...) creates a cheap, unloaded reference. Local bundle and
directory sources use URL through .bundle(url) and .directory(url);
Hugging Face uses .huggingFace("org/repo"). Async load() is the explicit
boundary for resolving, downloading, disk access, and runtime initialization.
Use loadSync() only from an existing worker thread.
Download Progress
load() blocks until the weights are on disk, so it has no progress to report
while it runs. Start the download separately and watch it:
let loader = Xybrid.model("qwen3-0.6b")
if let download = loader.download() {
for await status in download.progress() {
bar.progress = Float(status.progress)
label.text = "\(status.downloadedBytes) / \(status.totalBytes ?? 0)"
}
}
let model = try await loader.load() // cached: returns at onceprogress spans every file the model needs — weights plus companions such as a
vision projector — and never moves backwards, so the bar neither restarts per
file nor rewinds when a download retries. totalBytes is null when the source
publishes no size (a Hugging Face repo); downloadedBytes is exact either way,
which is what you need for megabytes, speed and time remaining.
Cancelling the consuming task only unsubscribes from updates. Call cancel()
to stop the transfer itself — it takes effect within one chunk read and
discards the partial file.
Input Envelopes
Audio (Speech Recognition)
let envelope = Envelope.audio(pcmData: audioData, sampleRate: 16000, channels: 1)
let result = try model.run(envelope: envelope)
print("Transcription: \(result.text ?? "")")Text (Text-to-Speech)
// Simple text
let envelope = Envelope.text("Hello, how are you?")
// With voice and speed
let envelope = Envelope.text("Hello", voice: "af_heart", speed: 1.0)
let result = try model.run(envelope: envelope)
if let audioBytes = result.audioBytes {
// Play or save audio
}Embedding
let envelope = Envelope.embedding(data: [0.1, 0.2, 0.3])
let result = try model.run(envelope: envelope)
if let vector = result.embedding {
// Process embedding vector
}Live Transcription
run() transcribes a finished audio buffer. A session transcribes speech as
it arrives, which is what dictation and live captioning need.
Audio must be PCM float32, mono, 16 kHz — converting from the platform's microphone format is the caller's job.
let model = try await Xybrid.model("whisper-tiny").load()
let session = try model.stream(config: .voiceActivity(modelDir: vadModelDir, language: "en"))
Task {
for await partial in session.partials() {
label.text = partial.text
}
}
// From the AVAudioEngine tap, converted to 16 kHz mono Float32:
try session.feed(samples: pcm)
// When the user stops talking:
let transcript = try session.flush()Partial text is cumulative, not a delta: render each one in place of the previous rather than appending. A partial produced before you subscribe is delivered immediately, so the opening words of an utterance are never lost.
flush() ends the session and returns the full transcript — and reports an
error if any chunk failed to transcribe, because that audio is gone and the
transcript would otherwise look complete. cancel() ends the session and
discards queued audio.
Voice-activity detection cuts chunks on speech boundaries instead of a fixed
clock, which avoids the stutter you get when a window lands mid-word. It needs
a Silero VAD model directory containing a model.onnx; none ships with the
SDK, and without one the engine falls back to fixed windows.
Result Handling
let result = try model.run(envelope: envelope)
if result.success {
switch result.outputType {
case "text":
print("Text: \(result.text!)")
case "audio":
playAudio(result.audioBytes!)
case "embedding":
process(result.embedding!)
default:
break
}
print("Latency: \(result.latency)s") // TimeInterval in seconds
} else {
print("Error: \(result.error ?? "Unknown")")
}XybridResult Properties
| Property | Type | Description |
|---|---|---|
success | Bool | Whether inference succeeded |
error | String? | Error message if failed |
outputType | String | "text", "audio", or "embedding" |
text | String? | Text output (ASR, LLM) |
audioBytes | Data? | Audio output (TTS) |
embedding | [Float]? | Embedding vector |
latencyMs | UInt32 | Inference latency in ms |
isFailure | Bool | Convenience: !success |
latency | TimeInterval | Latency in seconds |
Error Handling
The SDK uses a Swift error enum for type-safe error handling:
do {
let model = try await Xybrid.model("kokoro-82m").load()
let result = try await model.runAsync(envelope: envelope)
} catch XybridError.ModelNotFound(let modelId) {
print("Model not found: \(modelId)")
} catch XybridError.InferenceFailed(let message) {
print("Inference failed: \(message)")
} catch XybridError.InvalidInput(let message) {
print("Invalid input: \(message)")
} catch XybridError.IoError(let message) {
print("I/O error: \(message)")
} catch {
print("Unexpected error: \(error.localizedDescription)")
}Type Aliases
The SDK provides short aliases for convenience:
| Alias | Full Type |
|---|---|
Model | XybridModel |
Envelope | XybridEnvelope |
Result | XybridResult |
VoiceInfo | XybridVoiceInfo |
GenerationConfig | XybridGenerationConfig |
XybridSDKError | XybridError |
ModelLoader is a cheap model reference. Construct it through
Xybrid.model(...), then call load() explicitly before inference.
Platform Support
| Platform | Status | Accelerators |
|---|---|---|
| iOS | Supported | Metal, CoreML, ANE |
| macOS (Apple Silicon) | Supported | Metal, CoreML, ANE |
| macOS (Intel) | Supported | CPU |