Xybrid

xybrid run

Execute a hybrid pipeline

Execute a hybrid pipeline on the current device.

Usage

xybrid run --config <path>
xybrid run --pipeline <name>
xybrid run --pipeline <name> --dry-run

Options

OptionShortDescription
--config-cLoad a pipeline YAML file directly
--pipeline-pLook up <name>.yml under examples/
--bundle-bRun a .xyb bundle file directly
--input-textText input for text-based models (LLM, TTS)
--input-audioPath to an audio file for audio-based models (ASR) — WAV, MP3, OGG, or FLAC
--targetForce model format (onnx, coreml, tflite); auto-detected if omitted
--dry-runSimulate routing without execution
--policyLoad a policy bundle (YAML/JSON) that constrains routing — see Hybrid Execution › Policies
--traceEnable execution tracing with flame graph + LLM metrics output
--trace-exportExport the trace to a JSON file (Chrome trace format)

Behavior

  1. Parses YAML into PipelineConfig
  2. Runs policy → routing → execution loop
  3. Prints stage latency and routing summaries

In dry-run mode the CLI decides the first stage (policy and route) exactly as a real run would, at the current device snapshot, without running inference. Later stages are reported as unknown: their input is the previous stage's output, which does not exist without execution.

Examples

Run from config file

xybrid run --config ./pipelines/voice-assistant.yaml

Run named pipeline

xybrid run --pipeline hiiipe

Looks for hiiipe.yml or hiiipe.yaml in:

  • xybrid-cli/examples/
  • ./examples/

Dry run (simulate only)

xybrid run --pipeline hiiipe --dry-run

Output shows the first-stage decision without running inference:

  Stage 1 llm
  Declared         target: auto
  Provider         openai
  Policy           ALLOWED
  Reason           all policy checks passed
  Routing          local (default_local)

  Stage 2 tts
  Policy           UNKNOWN
  Routing          UNKNOWN
  Output           upstream output unavailable without execution

With a policy and Xybrid Platform

The primary hybrid example keeps the local model on the device and uses Xybrid Platform for cloud inference. Only the Xybrid API key belongs on the device; the upstream provider credentials are configured on the platform.

XYBRID_API_KEY="<your-xybrid-api-key>" xybrid run \
  -c crates/xybrid-cli/examples/hybrid-platform.yaml \
  --policy crates/xybrid-cli/examples/policies/prefer-cloud.yaml \
  --input-text "Summarise the plot of Dune in two sentences." -v

Swap in policies/local-only.yaml to keep the same prompt on the device, or policies/privacy.yaml to route by content. -v also prints the raw routing_decision JSON line.

Use --platform-url <base-url> for a staging or self-hosted platform. The example uses gpt-4o-mini, which the platform must have configured upstream.

Direct provider testing

crates/xybrid-cli/examples/hybrid-deepseek.yaml uses backend: direct and DEEPSEEK_API_KEY for direct DeepSeek calls. Keep it for manual provider checks; automated policy-routing tests use a loopback fake DeepSeek endpoint with dummy credentials.

Sample Output

  Stage 1 llm
  Routing          cloud
  Reason           [local] policy_route_cloud: policy rule 'route_cloud_if_0' prefers cloud: true (confidence: 90%)
  Backend          cloud:openai:gateway
  Time             842ms
  Output           Text

    Dune follows Paul Atreides ...

Pipeline Format

The run command expects YAML pipelines:

name: "voice-assistant"
stages:
  - whisper-tiny@1.0
  - target: integration
    provider: openai
    model: gpt-4o-mini
  - kokoro-82m@0.1

Tracing LLM Inference

When --trace is passed and the pipeline includes an LLM stage, the CLI prints an additional metrics block after the output:

📊 LLM Trace:
------------------------------------------------------------
  TTFT                  : 412 ms
  Decode TPS            : 42.80 tok/s
  Prefill TPS           : 184.20 tok/s
  Wallclock TPS         : 38.10 tok/s
  Mean ITL              : 14.30 ms
  p95 ITL               : 28 ms
  Emitted chunks        : 64
  Tokens generated      : 64
  Finish reason         : stop
------------------------------------------------------------

These values are read from the output envelope's metadata, populated by the LLM adapter's streaming path. Non-LLM stages (ASR, TTS, embedding) do not emit these fields and the block is skipped.

On this page