Canonical bytes
The same value always encodes to the same bytes, so a payload can be a cache key.
Already have a Zod, Valibot, or ArkType schema? Then you already have a binary
format. Install shorn and pass that schema to encode and
decode. There is no schema file to write, no code to generate, and
nothing to migrate.
// The validator you already use.const Person = z.object({ name: z.string(), age: z.int().nonnegative(), role: z.enum(["viewer", "editor", "admin"]),});
// It is also the wire format.const bytes = encode(Person, person);// → 8 canonical bytes, validated before encoding
const back = decode(Person, bytes);// → { name: string; age: number; role: "viewer" | "editor" | "admin" }npm install @chichurita/shorn JSON · 40 bytes
shorn · 8 bytes
Your schema already knows the field names, their order, and the enum members, so shorn sends only the values.
How it works walks through the eight bytes one at a time.
shorn reads the wire layout from your schema. Zod, Valibot, or ArkType still owns validation, transformations, and error messages, exactly as before.
encode(schema, value) and decode(schema, bytes)
with the validator you already have.
.proto file, then either generate code from it or set up
runtime reflection.
| shorn | 8 bytes |
|---|---|
| Avro (own schema) | 8 bytes |
| Protobuf (.proto + codegen) | 11 bytes |
| SchemaPack (own builder) | 13 bytes |
| msgpackr plain (field names) | 30 bytes |
| JSON (field names + text) | 40 bytes |
Pick a validator, then edit its schema and the payload. The real encoder runs in your browser tab, and nothing you type leaves it.
Suggestions appear as you type: arrows move, Enter or Tab picks. Otherwise Tab indents, and Escape first lets it leave the box.
JSON
shorn, as hex
Showing the last input that parsed.
loading the encoder…
The record above is a deliberate best case: booleans and enum members are what shorn shrinks hardest. Every figure quoted elsewhere on this page is measured on the benchmark fixtures instead, which are ordinary records.
For caches, RPC, worker messages, job queues, and storage, where both ends share the schema.
Not for readers in other languages, self-describing payloads, or schemas that
change independently. fingerprinted() rejects bytes written by an
older shape.
The same value always encodes to the same bytes, so a payload can be a cache key.
Every length is checked before allocating. Invalid UTF-8 and leftover bytes are rejected.
Small payloads normally cost CPU, because a compressor has to produce them. These come from the schema instead, in the same pass that encodes the value.
Person { age, name, sex }Event { active, actor: Person, id, metrics: { cpu, memory }, tags[], timestamp }Four fixtures: a Person, the same Person with non-ASCII text, one Event, and a batch of 100 Events in an array.