xybrid run
Execute a hybrid pipeline
Execute a hybrid pipeline on the current device.
Usage
xybrid run --config <path>
xybrid run --pipeline <name>
xybrid run --pipeline <name> --dry-runOptions
| Option | Short | Description |
|---|---|---|
--config | -c | Load a pipeline YAML file directly |
--pipeline | -p | Look up <name>.yml under examples/ |
--bundle | -b | Run a .xyb bundle file directly |
--input-text | Text input for text-based models (LLM, TTS) | |
--input-audio | Path to an audio file for audio-based models (ASR) — WAV, MP3, OGG, or FLAC | |
--target | Force model format (onnx, coreml, tflite); auto-detected if omitted | |
--dry-run | Simulate routing without execution | |
--policy | Load a policy bundle (YAML/JSON) that constrains routing — see Hybrid Execution › Policies | |
--trace | Enable execution tracing with flame graph + LLM metrics output | |
--trace-export | Export the trace to a JSON file (Chrome trace format) |
Behavior
- Parses YAML into
PipelineConfig - Runs policy → routing → execution loop
- Prints stage latency and routing summaries
In dry-run mode the CLI decides the first stage (policy and route) exactly as a real run would, at the current device snapshot, without running inference. Later stages are reported as unknown: their input is the previous stage's output, which does not exist without execution.
Examples
Run from config file
xybrid run --config ./pipelines/voice-assistant.yamlRun named pipeline
xybrid run --pipeline hiiipeLooks for hiiipe.yml or hiiipe.yaml in:
xybrid-cli/examples/./examples/
Dry run (simulate only)
xybrid run --pipeline hiiipe --dry-runOutput shows the first-stage decision without running inference:
Stage 1 llm
Declared target: auto
Provider openai
Policy ALLOWED
Reason all policy checks passed
Routing local (default_local)
Stage 2 tts
Policy UNKNOWN
Routing UNKNOWN
Output upstream output unavailable without executionWith a policy and Xybrid Platform
The primary hybrid example keeps the local model on the device and uses Xybrid Platform for cloud inference. Only the Xybrid API key belongs on the device; the upstream provider credentials are configured on the platform.
XYBRID_API_KEY="<your-xybrid-api-key>" xybrid run \
-c crates/xybrid-cli/examples/hybrid-platform.yaml \
--policy crates/xybrid-cli/examples/policies/prefer-cloud.yaml \
--input-text "Summarise the plot of Dune in two sentences." -vSwap in policies/local-only.yaml to keep the same prompt on the device, or
policies/privacy.yaml to route by content. -v also prints the raw
routing_decision JSON line.
Use --platform-url <base-url> for a staging or self-hosted platform. The
example uses gpt-4o-mini, which the platform must have configured upstream.
Direct provider testing
crates/xybrid-cli/examples/hybrid-deepseek.yaml uses backend: direct and
DEEPSEEK_API_KEY for direct DeepSeek calls. Keep it for manual provider checks;
automated policy-routing tests use a loopback fake DeepSeek endpoint with dummy
credentials.
Sample Output
Stage 1 llm
Routing cloud
Reason [local] policy_route_cloud: policy rule 'route_cloud_if_0' prefers cloud: true (confidence: 90%)
Backend cloud:openai:gateway
Time 842ms
Output Text
Dune follows Paul Atreides ...Pipeline Format
The run command expects YAML pipelines:
name: "voice-assistant"
stages:
- whisper-tiny@1.0
- target: integration
provider: openai
model: gpt-4o-mini
- kokoro-82m@0.1Tracing LLM Inference
When --trace is passed and the pipeline includes an LLM stage, the CLI
prints an additional metrics block after the output:
📊 LLM Trace:
------------------------------------------------------------
TTFT : 412 ms
Decode TPS : 42.80 tok/s
Prefill TPS : 184.20 tok/s
Wallclock TPS : 38.10 tok/s
Mean ITL : 14.30 ms
p95 ITL : 28 ms
Emitted chunks : 64
Tokens generated : 64
Finish reason : stop
------------------------------------------------------------These values are read from the output envelope's metadata, populated by the LLM adapter's streaming path. Non-LLM stages (ASR, TTS, embedding) do not emit these fields and the block is skipped.