Skip to content

Large Payloads

Arrow batches can get big. vgi-rpc gives you three cooperating mechanisms for keeping large payloads off the inline wire:

  • Response size caps — externalize a batch when configured, or refuse any response that still exceeds its byte budget.
  • External-location offloading — replace an oversized batch with a tiny zero-row “pointer” batch that the peer resolves out-of-band from object storage.
  • Request upload URLs — let a client upload a large request payload to a pre-signed URL and send the server a pointer instead.

All of these are most relevant to the HTTP transport. External-location resolution also works on the pipe and subprocess client transports.

Two HttpHandlerOptions cap how much a single response may produce:

Option Applies to Behavior
maxResponseBytes unary, stream-exchange, and each producer turn Always hard: overshoot replaces the response with an EXCEPTION batch and never returns a continuation cursor.
maxExternalizedResponseBytes every response that externalizes Always hard — externalized uploads have no continuation-token escape valve.
import { createHttpHandler } from "@query-farm/vgi-rpc";
const handler = createHttpHandler(protocol, {
maxResponseBytes: 5_000_000, // 5 MB inline body cap
maxExternalizedResponseBytes: 50_000_000, // 50 MB external-upload cap per response
});

When a hard cap is exceeded the handler discards the data it had built and returns a stream carrying only an EXCEPTION batch (surfaced to the client as an RpcError). maxResponseBytes applies to unary, exchange, and every producer turn. An oversized producer turn does not carry a cursor. maxExternalizedResponseBytes is pre-flighted before an upload is incurred so rejected bytes never leave the server.

Producer dispatch remains lock-step: one HTTP request invokes the producer exactly once. If an invocation stays within its limit and leaves the stream unfinished, the response carries a zero-row continuation-token batch and the client resumes by calling /{protocol}/{method}/exchange with that token.

The handler advertises these caps to clients via response headers VGI-Max-Response-Bytes and VGI-Max-Externalized-Response-Bytes. Clients first issue a cached OPTIONS /health, require exactly VGI-Accept-Max-Response-Bytes-Support: true, and send VGI-Accept-Max-Response-Bytes on every RPC. They also require the exact support header on every RPC response and locally enforce the minimum of their accepted maximum and the advertised server cap while incrementally reading the decoded Fetch stream. Browser/workerd clients default to 64 MiB; native clients default to 256 MiB. Configured and accepted response limits have a 64 KiB minimum. Undefined server limits remain unbounded, subject to the client’s limit.

Instead of refusing a large batch, the server can upload it to object storage and leave only a pointer batch on the wire: a zero-row batch (same schema) whose custom metadata carries vgi_rpc.location (the retrieval URL) and vgi_rpc.location.sha256 (the SHA-256 digest of the raw IPC bytes). The peer detects the pointer, fetches the data, verifies the checksum, and continues as if the batch had arrived inline. See the wire-protocol reference for the exact pointer-batch format.

Externalized payloads do not count toward maxResponseBytes — only the tiny pointer batch rides on the wire.

Pass an ExternalLocationConfig as the externalLocation option:

import { createHttpHandler, type ExternalStorage, type ExternalLocationConfig } from "@query-farm/vgi-rpc";
class S3Storage implements ExternalStorage {
async upload(data: Uint8Array, contentEncoding: string): Promise<string> {
// Persist `data` (Arrow IPC bytes, possibly zstd-compressed) and return
// an HTTPS URL the peer can GET. `contentEncoding` is "zstd" when the
// config enabled compression, otherwise "".
return await putObjectAndSign(data, contentEncoding);
}
}
const externalLocation: ExternalLocationConfig = {
storage: new S3Storage(),
externalizeThresholdBytes: 1_048_576, // default 1 MB; batches at/above this are offloaded
compression: { algorithm: "zstd", level: 3 }, // optional; omit to upload uncompressed
// urlValidator defaults to httpsOnlyValidator; pass null to disable validation
};
const handler = createHttpHandler(protocol, {
externalLocation,
maxExternalizedResponseBytes: 50_000_000,
});

The handler advertises whether externalization is enabled via the VGI-Externalization-Enabled response header.

Field Type Description
storage ExternalStorage Backend used to upload batch IPC bytes.
externalizeThresholdBytes? number Minimum batch IPC byte size that triggers offloading. Default: 1048576 (1 MB).
compression? { algorithm: "zstd"; level?: number } Optionally zstd-compress uploaded data. level defaults to 3.
urlValidator? ((url: string) => void) | null Called before fetching a pointer URL; throw to reject. Default: httpsOnlyValidator. Pass null to disable validation entirely.

The storage backend is a single-method interface you implement:

interface ExternalStorage {
/** Upload IPC data and return a URL for retrieval. */
upload(data: Uint8Array, contentEncoding: string): Promise<string>;
}

data is the serialized Arrow IPC stream for the batch (zstd-compressed when compression is configured). contentEncoding is "zstd" in that case, otherwise "" — if you persist it as the object’s Content-Encoding, resolveExternalLocation will transparently decompress on the read side. The returned URL must be fetchable by the peer (and, by default, must be HTTPS). Object lifecycle/cleanup is your responsibility — vgi-rpc never deletes uploaded objects.

The default URL validator. It throws unless the URL uses the https: scheme:

import { httpsOnlyValidator } from "@query-farm/vgi-rpc";
httpsOnlyValidator("https://bucket.example/abc"); // ok
httpsOnlyValidator("http://bucket.example/abc"); // throws

Supply your own urlValidator (e.g. an allowlist of trusted hosts) to harden against fetching from attacker-controlled locations, or set it to null to skip validation (only for trusted, e.g. loopback, deployments).

These helpers underlie the automatic offloading and are exported for advanced/manual use:

import {
maybeExternalizeBatch,
resolveExternalLocation,
makeExternalLocationBatch,
isExternalLocationBatch,
} from "@query-farm/vgi-rpc";
  • maybeExternalizeBatch(batch, config?) — write path. If config.storage is set, the batch has rows, and its IPC size is at or above the threshold, it serializes the batch (optionally zstd-compressing), uploads it, and returns a pointer batch carrying the location and SHA-256. Otherwise returns the batch unchanged.
  • resolveExternalLocation(batch, config?) — read path. If the batch is a pointer (and config is provided), it validates the URL, fetches it, decompresses if the response is Content-Encoding: zstd (capped at 16× the compressed size as a decompression-bomb defense), verifies the SHA-256, and returns the resolved data batch. Non-pointer batches pass through unchanged.
  • makeExternalLocationBatch(schema, url, sha256?) — builds a zero-row pointer batch with the given schema, setting vgi_rpc.location (and vgi_rpc.location.sha256 when provided).
  • isExternalLocationBatch(batch) — returns true for a zero-row batch carrying vgi_rpc.location (and not a log/error batch).

Per-call offloading serializes, compresses, and uploads a new object on every call. A result that is large and changes rarely — a whole catalog, identical for every caller until redeploy — can instead be published once and answered with the same pointer on every later call.

A unary handler may return an ExternalRef in place of its result values. The dispatcher then writes the pointer batch for that ref directly:

  • the result is not built or validated, nothing is serialized, compressed or uploaded during the call;
  • it is used whether or not the server has externalLocation configured, and regardless of externalizeThresholdBytes — a ref is never inlined, however small;
  • it does not count toward maxExternalizedResponseBytes (nothing is uploaded), but the tiny pointer still goes through the maxResponseBytes wire budget;
  • the same on every transport (stdio, unix, tcp, HTTP). Unary methods only.

publishExternal(batch, storage, { compression?, includeSha256? }) serializes a 1-row result batch exactly as the per-call externalizer does, hashes the raw IPC bytes, compresses when asked (pass your ExternalLocationConfig.compression to match it), uploads once, and returns the ExternalRef. publishExternalResult(schema, values, storage, options?) builds that batch from the same { result: value } record a handler would return.

import { ExternalRef, Protocol, publishExternalResult, str } from "@query-farm/vgi-rpc";
const protocol = new Protocol("Catalog");
let catalogRef: Promise<ExternalRef> | undefined;
protocol.unary("catalog", {
params: {},
result: { result: str },
handler: () => {
// Cache the promise, so concurrent first calls still upload once.
catalogRef ??= publishExternalResult(
protocol.getMethod("catalog")!.resultSchema,
{ result: buildCatalog() },
storage,
{ compression: { algorithm: "zstd", level: 3 } },
);
return catalogRef;
},
});

You can also construct new ExternalRef(url, sha256?) by hand for an object published out of band. The object must be an Arrow IPC stream (optionally Content-Encoding: zstd) whose schema is the method’s result schema and which holds exactly one 1-row batch. sha256, when given, must be 64 lowercase hex characters over the raw (pre-compression) bytes; omit it (or pass includeSha256: false to publishExternal) and the pointer carries no vgi_rpc.location.sha256, so clients skip the content check — useful for an object rewritten in place or too large to hash. An empty url or malformed digest throws.

Clients need no change: a pointer is a pointer. The caller owns caching the ref and the object’s lifecycle — a long-lived ref must not point at an object under the short-TTL lifecycle rule used for per-call uploads, and a pre-signed URL expires, so re-sign or rebuild the ref before then. Only return a ref to callers who are all entitled to the same content.

The size caps and external storage above protect the response side. For large requests, the client can upload its payload to a pre-signed URL and send the server a pointer instead of the inline body.

Set uploadUrlProvider to expose POST {prefix}/__upload_url__/init. The route is exempt from maxRequestBytes (it exists precisely to escape that limit). A client POSTs a tiny request asking for count URL pairs (clamped to 100); the handler responds with an Arrow batch of upload_url, download_url, and expires_at rows.

import { createHttpHandler } from "@query-farm/vgi-rpc";
const handler = createHttpHandler(protocol, {
maxRequestBytes: 10_000_000, // inline request bodies above this should externalize
maxUploadBytes: 500_000_000, // advertised max external upload (VGI-Max-Upload-Bytes)
uploadUrlProvider: {
async generateUploadUrl() {
const { putUrl, getUrl, expires } = await signUploadPair();
return { uploadUrl: putUrl, downloadUrl: getUrl, expiresAt: expires };
},
},
// To then resolve the uploaded request payload server-side, also configure
// externalLocation with a storage/validator.
externalLocation,
});

The provider implements generateUploadUrl(), returning { uploadUrl, downloadUrl, expiresAt }: uploadUrl is the pre-signed PUT the client uploads to, downloadUrl is the GET the server fetches from, and expiresAt is the pair’s UTC expiry. Implementations must be safe to call concurrently, and the operator owns object cleanup.

When configured, the handler advertises VGI-Upload-URL-Support: true and (if set) VGI-Max-Upload-Bytes on responses. On the dispatch side, when a request arrives as an external-location pointer, the unary handler resolves it via the configured externalLocation, re-attaches the outer dispatch metadata (method name, version, request id), and parses parameters as usual.

When the client (httpConnect) sees that the server advertises upload-URL support and a maxRequestBytes smaller than the body it is about to send, it transparently fetches a pre-signed pair from /__upload_url__/init, PUTs the request IPC bytes to uploadUrl, and sends the server a pointer referencing downloadUrl. It also retries this way if a plain POST returns 413 Payload Too Large. The client passes its externalLocation.urlValidator through to validate vended URLs.

The client connect functions — httpConnect, pipeConnect, and subprocessConnect — accept an externalLocation option of the same ExternalLocationConfig type. The client uses it to:

  • Resolve externalized response batches it receives (fetch + verify + decompress the pointer).
  • Source the urlValidator used when externalizing large requests (HTTP only).
import { httpConnect, httpsOnlyValidator, type ExternalStorage } from "@query-farm/vgi-rpc";
const client = httpConnect("https://api.example.com", {
externalLocation: {
storage: myStorage, // ExternalStorage (only used for client-vended request uploads, if any)
urlValidator: httpsOnlyValidator,
},
});