Skip to content
Query-farmPublic

About

Rust port of the VGI worker SDK + example fixtures (DuckDB VGI protocol)

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

Β 

History

381 Commits

Folders and files

Repository files navigation

Vector Gateway Interface logo

VGI for Rust

Add your own functions and tables to DuckDB with Rust and Apache Arrow.
Built by 🚜 Query.Farm

CI crates.io crates.io downloads docs.rs License

A VGI worker is a small Rust program that DuckDB talks to over Apache Arrow IPC. It can expose scalar / table / aggregate functions and whole catalogs (schemas, tables, views) that behave like native DuckDB objects. DuckDB launches your worker for you when a query needs it β€” you never run a server by hand.

This repo is the Rust worker SDK (vgi). It is byte-for-byte wire-compatible with the canonical Python SDK, so a Rust worker drops in behind the same ATTACH ... (TYPE vgi). Built on vgi-rpc; stock arrow-rs 59.x, MSRV 1.97.

HTTP producer batching and worker-visible response budgets are covered in docs/response-budgets.md.

Why a worker instead of a C++ extension?

Traditional DuckDB extension VGI worker
Written in C/C++, compiled and linked against DuckDB Written in Rust, one standalone binary
Must be rebuilt for each DuckDB version Version independent
Complex build / signing / release cycle cargo build, ship the binary
Runs in-process Process isolation

Reach for it when you want to: call REST APIs from SQL, run ML inference, expose an external database / API / filesystem as a queryable catalog, or ship domain-specific functions to your team as a single binary.

Your first worker

1. Create a project and add the dependencies:

# Cargo.toml
[dependencies]
vgi = "0.35"
vgi-rpc = "0.25"
arrow-array = "59"
arrow-schema = "59"

2. Write a function and serve it:

// src/main.rs
use std::sync::Arc;

use arrow_array::{cast::AsArray, ArrayRef, RecordBatch, StringArray};
use arrow_schema::DataType;
use vgi::{ArgumentMonotonicity, ArgSpec, FunctionMetadata, ProcessParams, ScalarFunction, Worker};
use vgi_rpc::{Result, RpcError};

/// `upper_case(s)` β€” uppercase a string column.
struct UpperCase;

impl ScalarFunction for UpperCase {
    fn name(&self) -> &str {
        "upper_case"
    }

    fn metadata(&self) -> FunctionMetadata {
        FunctionMetadata {
            description: "Convert string values to uppercase".into(),
            return_type: Some(DataType::Utf8),
            argument_monotonicity: Some(vec![ArgumentMonotonicity::Unknown]),
            ..Default::default()
        }
    }

    fn argument_specs(&self) -> Vec<ArgSpec> {
        vec![ArgSpec::column("value", 0, "varchar", "String to uppercase")]
    }

    fn process(&self, params: &ProcessParams, batch: &RecordBatch) -> Result<RecordBatch> {
        let col = batch.column(0).as_string::<i32>();
        let upper: StringArray = col.iter().map(|v| v.map(str::to_uppercase)).collect();
        let out: ArrayRef = Arc::new(upper);
        RecordBatch::try_new(params.output_schema.clone(), vec![out])
            .map_err(|e| RpcError::runtime_error(e.to_string()))
    }
}

fn main() {
    let mut worker = Worker::new();
    worker.register_scalar(UpperCase);
    worker.run(); // serves stdio (default), --unix <path>, or --http
}

argument_monotonicity is optional and scalar-only. A present vector has one entry per ordered argument declaration. Fixed, defaulted, and constant arguments each occupy one slot; a vararg declaration occupies one slot even when a call expands it. Named call syntax does not change this order.

3. Build it (cargo build --release), then call it from a DuckDB engine that has the vgi extension. The vgi extension currently ships with Query Farm's Haybarn DuckDB distribution, which starts with no install via uvx haybarn-cli. From your project directory:

-- Haybarn ships the `vgi` extension. DuckDB LAUNCHES the worker for you;
-- LOCATION is the command it runs, and the alias 'demo' is what you
-- qualify functions with in SQL.
ATTACH 'demo' (TYPE vgi, LOCATION './target/release/my-worker');

SELECT demo.main.upper_case(name) FROM (VALUES ('alice'), ('bob')) t(name);
-- ALICE
-- BOB

LOCATION gotcha: the path is resolved relative to DuckDB's working directory, not your project. If the worker isn't found, use an absolute path.

Iterating

Change your Rust, rebuild, and re-attach β€” DuckDB pools the worker per attachment, so DETACH demo; ATTACH 'demo' (...) (or a fresh session) picks up a new build.

Troubleshooting

  • ATTACH can't find the worker β€” LOCATION is relative to DuckDB's working directory; use an absolute path.
  • Catalog Error: ... does not exist β€” qualify with the attach alias (demo.main.upper_case) or run USE demo;.
  • Runtime / type errors β€” errors returned from process (and bind-time argument_specs type checks) surface directly in DuckDB's error message.
  • Telling bad input from a bug β€” every error carries a gRPC-style code (vgi_rpc.error_code). The SDK codes its own: bad arguments and type-bound failures are INVALID_ARGUMENT, unknown functions / tables / schemas are NOT_FOUND, writes to a read-only catalog are FAILED_PRECONDITION, unsupported operations are UNIMPLEMENTED; anything else is UNKNOWN. Use vgi::errors::{invalid_argument, not_found, failed_precondition, unimplemented} to code your own worker's errors the same way.

Function types

Type Trait SQL pattern Use case
Scalar ScalarFunction SELECT f(col) FROM t Per-row transforms (1:1)
Table TableFunction SELECT * FROM f(args) Generate / scan data
Table-In-Out TableInOutFunction SELECT * FROM f((SELECT …)) Streaming transforms
Table-Buffering TableBufferingFunction SELECT * FROM f((SELECT …)) Aggregate-then-emit (sink β†’ combine β†’ source)
Aggregate AggregateFunction SELECT f(col) … GROUP BY … Grouped / window / streaming aggregates

Each trait is small: name, metadata, argument_specs, an on_bind to resolve the output schema, and process (or the buffering / aggregate lifecycle methods). Projection & filter pushdown, ORDER BY / TABLESAMPLE hints, settings, secrets, bearer auth, and a cross-process state store are handled for you.

Beyond functions: full catalogs

Worker::set_catalog exposes a complete catalog β€” schemas, function-backed tables, views, and macros β€” with constraints, column statistics, time travel (AT), and secondary catalogs attachable by name:

ATTACH 'external_db' (TYPE vgi, LOCATION './my-catalog-worker');

SELECT * FROM external_db.main.users;            -- a function-backed table
SELECT * FROM external_db.analytics.daily_view;  -- a view
SELECT external_db.main.transform(col) FROM t;   -- a function

Transports

Worker::run picks the transport from argv: stdio (default), Unix socket (--unix <path>, the launcher contract), or HTTP (--http, Arrow-IPC over HTTP with AEAD-sealed stateless stream tokens and optional bearer auth).

Every transport hosts vgi.v2, vgi_rpc.Reflection.v1, and any protocols the worker adds with Worker::hosted_protocols; HTTP additionally hosts vgi_rpc.Identity.v1 when the worker sets Worker::resolve_token and/or Worker::mint_grant; with a grant key (--grant-key / VGI_RPC_GRANT_KEYS) it mints sealed grants and accepts them back as bearer credentials. With a grant source and a signing key (VGI_SIGNING_KEY), HTTP also hosts vgi.attach_tickets.v1: seal_attach turns a user's ATTACH (secret options included) into a vgia1. ticket that a runner holding the user's grant redeems with the single attach option vgi_attach_ticket. See docs/hosted-protocols.md.

Over HTTP the worker seals the attach_opaque_data and transaction_opaque_data values it hands the client (XChaCha20-Poly1305 under VGI_SIGNING_KEY / Worker::signing_key, or a key generated at startup when neither is set), binds each to its caller and each transaction to its attach, and refuses anything that does not open with the one error <field> not recognized. Secret attach options never enter either value on any transport. See vgi::opaque.

Protocol overview

VGI uses vgi-rpc, an Apache-Arrow-IPC RPC framework, for all DuckDB ↔ worker communication. You don't write to this directly β€” the traits handle it β€” but here's what happens per query:

DuckDB (client)                      VGI worker
  │──── bind(request) ─────────────▢ β”‚  function name, args, input schema
  │◀─── BindResponse ───────────────  β”‚  output schema (your on_bind)
  │──── init(request) ─────────────▢ β”‚  start the processing stream
  │◀─── stream header ──────────────  β”‚  execution_id, max_workers
  │──── process(batch) ────────────▢ β”‚
  │◀─── output batch ───────────────  β”‚  your process(batch)
  │──── [stream close] ────────────▢ β”‚

Workspace layout

crate published summary
vgi/ βœ… crates.io Β· docs.rs The worker SDK: function models, declarative catalogs, wire dispatch, transports.
vgi-example-worker/ β€” A fixture worker registering every function kind and full catalogs; drives the integration suite. publish = false.

Read vgi-example-worker/src/ for a working example of every trait β€” scalar, table, table-in-out, buffering, aggregate, and catalog-backed tables/views.

Testing your own worker

The fastest check is to call your function from a DuckDB session (see "Your first worker"). For automated tests, drive the worker from Rust with vgi-rpc's client, or shell out to a DuckDB session from your test harness.

Running the SDK's integration suite (contributors)

The full behavioral suite is the canonical VGI C++ integration suite (test/sql/integration/* in the vgi extension repo), which drives DuckDB's unittest binary against the example worker. It passes across all three transports (8176 assertions on subprocess, 7774 on HTTP, 0 failures):

cargo build --release
scripts/run_tests.sh            # subprocess transport, full in-scope suite
LAUNCH=1 scripts/run_tests.sh   # launcher (Unix socket) transport
scripts/run_http_tests.sh       # HTTP transport

cargo fmt / clippy / build / doc run in CI.

For native iroh:// / httpi:// clients and bridge-ready workers, see Iroh operations.

Development

vgi depends on the published vgi-rpc from crates.io. To develop against an unreleased vgi-rpc checkout, add an uncommitted patch to the root Cargo.toml:

[patch.crates-io]
vgi-rpc = { path = "../vgi-rpc-rust/vgi-rpc" }
cargo build --workspace
cargo clippy --workspace --all-targets --all-features -- -D warnings
# every feature of vgi / vgi-client on its own (cargo install cargo-hack)
cargo hack clippy -p vgi -p vgi-client -p vgi-example-worker --each-feature --all-targets -- -D warnings
cargo test --doc -p vgi
cargo fmt --all

License

Query Farm Source-Available License v1.0 β€” see LICENSE. Copyright Β© 2025, 2026 Query Farm LLC.

About

Rust port of the VGI worker SDK + example fixtures (DuckDB VGI protocol)

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages