Performance
How to read the ratings
Section titled “How to read the ratings”These ratings are a rough guide to the cost of one call. A Slow entry can still be 2–4× faster than
postMessage, depending on the payload and the work being done.
- Best — almost no per-call overhead
- Fast — inexpensive for most workloads
- Good — suitable for normal workloads
- Fair — fine occasionally; watch hot paths
- Slow — consider another representation in tight loops
Thresholds
Section titled “Thresholds”The thresholds below are intentionally broad. They describe the approximate cost of one call, not the total time spent running the task:
- Best: < 1 µs
- Fast: < 2 µs
- Good: < 4 µs
- Fair: < 9 µs
- Slow: > 9 µs
Benchmark context
Section titled “Benchmark context”The figures on this page were measured on an Apple M3 Ultra running Node 24.12.0 on arm64-darwin, at roughly 3.86 GHz. Treat them as useful comparisons, not promises: your runtime, CPU, payload shape, and task itself will change the result.
How payloads move
Section titled “How payloads move”Most performance differences come down to how much data has to cross the worker boundary and whether that data is copied.
- Small values. Numbers, booleans, short strings, and similar values fit in the call header, so encoding and decoding are very cheap.
- Small binary payloads. Typed arrays and other small values use the transport’s preallocated space and usually stay fast.
- Larger payloads. The transport may need to allocate more space and copy the data. This is still efficient, but the cost becomes visible in a hot loop.
- Shared memory.
SharedArrayBufferandProcessSharedBufferavoid copying the bytes. Knitting passes a handle instead, so the cost of sending a 1 KiB buffer and a 64 MiB buffer is roughly the same.BufferReference(knitting/unsafe) is the thread-only move variant: it detaches the source and hands the bytes to the worker.
Typical cost by value type
Section titled “Typical cost by value type”These ratings describe one value passed in a single call. They are useful for choosing a representation, but measure your real workload before tuning around them.
| Value | Typical cost |
|---|---|
Primitives: boolean, undefined, null | Best |
Numbers: number | Best |
Time/IDs: Date | Best |
Strings: small string | Best |
Symbols: Symbol.for | Fast |
BigInt: small bigint | Best |
BigInt: large bigint | Fast |
| Binary: typed arrays | Best |
Views: DataView | Good |
| Structured: JSON object | Good |
| Structured: JSON array | Good |
Errors: Error | Slow |
Tuning the pool
Section titled “Tuning the pool”Thread count
Section titled “Thread count”Thread count depends on what the host thread must still do. For a mixed HTTP
service, start with one worker: the host still accepts connections, routes,
encodes requests, and sends responses. Add workers only when CPU-heavy calls
queue and lower tail latency is worth the extra coordination. For independent
batch compute, os.availableParallelism() - 1 is a reasonable first trial.
Every worker also consumes memory for payload buffers, shared locks, and cancellation state. Adding threads beyond the available cores usually stops helping. See Multi-threading for a server-focused selection process and timer configuration.
Native work stealing
Section titled “Native work stealing”Compatible multi-worker pools use native work stealing by default. The host publishes calls to one shared submit region, and workers claim available tasks as they become free. Responses still use private return lanes, so the task API and promise behavior do not change.
This helps most when many CPU-bound calls compete for workers and task durations
are uneven. It is unlikely to be the main lever for tasks that mostly wait on
databases, networks, or other external I/O. Set host.steal: false for a
private-lane baseline, or tune host.stealRegionLanes for the workload:
- Wider regions reduce arbitration overhead for many cheap, similarly sized calls.
- Narrower regions expose more independent work for expensive or uneven calls.
1is a useful starting point when one task can take much longer than another.
The host doorbell is a separate completion optimization. Node.js and Bun
thread pools can use Atomics.waitAsync instead of repeatedly polling for
responses; Deno, process, compiled, and browser pools use the polling fallback.
Compare like with like when benchmarking: keep the worker topology fixed and
change one of host.steal or host.doorbell at a time. See Work stealing
for the full option and compatibility details.
Inliner
Section titled “Inliner”The inliner runs eligible calls without sending them through a worker, skipping
transport encoding and decoding. It is most useful for very small tasks, such
as simple arithmetic. For example,
inliner: { position: "last", batchSize: 64 } can improve throughput. See the
Inliner guide.
Permissions
Section titled “Permissions”Strict permissions add a small amount of startup work for each worker, such as generating flags and resolving lock files. They do not add measurable overhead to individual calls once the workers are running.
Payload sizing
Section titled “Payload sizing”payload.payloadInitialBytes, payload.payloadMaxByteLength, and
payload.maxPayloadBytes control the transport buffer used by each worker.
Increasing the initial size avoids growth later, but uses more memory up front.
For consistently small payloads such as primitives and short strings, the
defaults—4 MiB initially, 64 MiB maximum length, and an 8 MiB dynamic
payload cap—are usually enough.
Choosing threads or processes
Section titled “Choosing threads or processes”Threads and processes have different memory boundaries, so the best zero-copy option depends on which one you use:
| Runtime | Isolation | Zero-copy tools |
|---|---|---|
thread (default) | shares the host address space | SharedArrayBuffer, BufferReference (move) |
process | separate memory and permissions | ProcessSharedBuffer (OS shared memory) |
A SharedArrayBuffer or BufferReference cannot cross a process boundary. Use
ProcessSharedBuffer when the worker runs in a separate process. See
Shared memory and
Buffer reference.
Because a process worker is a child process, you can start it through another
tool using worker.processCommandPrefix. This is useful for a sandbox such as
bwrap or a container such as Docker:
worker: { runtime: "process", processCommandPrefix: ["bwrap", "--unshare-all", "--ro-bind", "/", "/"],}See Process workers for the full wrapper recipes.
Keep request bodies off the main thread
Section titled “Keep request bodies off the main thread”call.*() accepts promises for supported inputs. In an HTTP handler, you can
pass the request body’s promise straight to a task instead of awaiting it on
the request thread:
app.post("/jwt", async (c) => { const responseJson = await handlers.call.issueJwt(c.req.arrayBuffer());
return c.body(responseJson ?? "Bad request", responseJson ? 200 : 400, { "content-type": "application/json; charset=utf-8", });});The promise itself is not faster. The benefit is that the request thread can hand the body to Knitting immediately:
c.req.arrayBuffer()already returns a promise, so forwarding it skips anawaitin the handler.- UTF-8 decoding and JSON parsing happen in the worker, not on the request thread.
ArrayBufferstays on the binary fast path.
Metadata and body with Envelope
Section titled “Metadata and body with Envelope”When a task needs both request metadata and the raw body, put them in an
Envelope. The header carries the metadata and the payload carries the bytes.
Use .then(...) to build the envelope from the body promise without awaiting
the body on the request thread:
import { Envelope } from "knitting";
app.post("/upload", async (c) => { const result = await handlers.call.storeUpload( c.req.arrayBuffer().then( (body) => new Envelope( { contentType: c.req.header("content-type") ?? "application/octet-stream" }, body, ), ), );
return c.json(result);});For a large binary body sent to a thread worker, use a BufferReference
instead of an ArrayBuffer. The bytes move without a copy; only the body line
changes:
import { BufferReference } from "knitting/unsafe";
new Envelope( { contentType: c.req.header("content-type") ?? "application/octet-stream" }, new BufferReference(body), // moves the body bytes without copying);This works best when the route is mainly forwarding data and the worker does the parsing, such as SSR or JWT issuance. If the main thread needs to inspect or validate the body first, await it there instead.
Choosing how to return data
Section titled “Choosing how to return data”Returning data has the same costs as sending it, just in the other direction. Choose the return type based on the size of the result and how the data is already represented:
- JSON object / array — serialized on the worker and parsed again on the host, so it makes two passes over the data. This is fine for small results but expensive for large ones.
SharedArrayBuffer/ProcessSharedBuffer— shared memory is usually the cheapest way to return bytes when the result can use it.ProcessSharedBufferalso works with process workers.BufferReference— useful for large binary results, especially when a library gives you an ordinaryArrayBufferthat you cannot turn into shared memory. Returned references can stay zero-copy on Node with native ownership; the safe default makes one copy on Deno, Bun, and some Node builds. See the Buffer reference guide for the explicit borrow option and its lifetime rules.
See Payloads and Buffer reference for the full type list.