Skip to content
Assay

Reference

Performance

Every row below came off one arm64 Mac: warm, -O, minimum of five rounds, on the corpus in the repository. None of them is a claim about your machine, or about another platform. Linux and x86-64 get their own sections in Benchmarks/RESULTS.md, and where a number flips across platforms this page says so rather than quoting the flattering half.

Arm Number Against
struct decode, full corpus 8.75× mean over 25 files (4.95–19.83) JSONDecoder
prefix decode + unknown-key skip 5.52× over 45 files (2.65–8.22) JSONDecoder
generic value model 3.31× over 75 files JSONSerialization
falsification arm (apimodel, 5 sizes) 5.21× mean (7.94× float-dense) JSONDecoder
vs ZippyJSON (simdjson + Codable) 3.20–3.66× faster, over four sizes ZippyJSON, which is 1.62–1.82× over Foundation here
vs yyjson, use-case shape 0.71× (loses) yyjson parse + extraction
vs yyjson, float-dense 0.71× (loses) same
vs yyjson, DOM vs DOM 0.16× (loses) yyjson_read
YAML node parse 8.28× Yams compose
YAML struct decode 17.47× Yams YAMLDecoder
XML tree parse 2.53× (macOS; 0.96× on Linux, 2026-08-18) Foundation XMLParser
TOML node parse 3.69× toml++ via TOMLKit
TOML struct decode 6.92× TOMLKit TOMLDecoder
Date fields 5.58× mean over 5 sizes JSONDecoder + .iso8601
binary plist 4.33× Foundation PropertyListDecoder
XML plist 1.29× Foundation PropertyListDecoder
union vs its variant 1.19× (the tag scan) the variant decoded directly
@Inline vs nesting 0.85× (faster) the nested @Schema it replaces
@Wraps vs @Validate 1.98× (slower) the plain field + rule it is sugar for
encoding, 50 / 200 items 8.28× / 8.02× JSONEncoder
cold start, 60 types 6.1× first decode (median); 5.2× steady JSONDecoder
multi-megabyte documents 8.41–9.02×, ~1,020 MB/s, flat JSONDecoder
total allocations, 50 items 159 against Foundation’s 377 JSONDecoder
T.validate(_:) 40 ns per value, 1 block; 47 ns/row batched, 0.11× a decode —
live allocations, apimodel-8k struct gated, PASS absolute thresholds
compile time, 10 fields 84.4 ms/type (budget 100) Codable: 4.03×

Assay does not need to beat simdjson. It needs to not have a KeyedDecodingContainer.

That is the whole thesis. ZippyJSON bolted simdjson — the fastest JSON parser in existence — onto Decodable and got 1.38× over Foundation. Apple’s own prototype, which changes nothing about parsing and only deletes the container protocol, reports about 6×. Roughly 83% of a Swift decode is the Codable boundary, and a macro deletes it at compile time.

Which is why Assay comes out 3.2–3.7× faster than ZippyJSON here. Scalar Swift, no SIMD anywhere, against simdjson underneath. The container is the only structural difference.

And ZippyJSON is not being hobbled: it measures 1.62–1.82× over Foundation on this machine, better than the 1.38× it was originally cited at.

It is not faster than C. Against yyjson, Assay loses: 0.71× on the use-case shape, 0.16× building a tree. Those rows are in the table above, published rather than omitted, because a benchmark page that lists only its wins is an advertisement.

Those two rows measure different things, and the gap between them is the interesting part.

The use-case arm makes yyjson do the whole job: parse, then walk the tree and pull the fields out. A document you still have to walk is not a decoded value. There Assay lands at about two thirds of hand-tuned C, which is a fine place to be.

The tree-against-tree arm is where it loses about 6×, and it compares two different products. yyjson writes one arena with its strings pointing into your input buffer. JSON.Value is a Swift enum tree of real Strings and Arrays you can store, compare and hash. Closing that gap means giving up all three.

The value-model path has no thesis behind it. JSON.Value measures about 3.3× over JSONSerialization — comfortably ahead of the thing you would otherwise reach for, and nowhere near the struct path, because building a tree has no Codable boundary to delete. Use @Schema when you know the shape; that is where the argument applies.

XML is a Darwin-only claim. Assay’s XML parser measures 2.53× over Foundation on macOS and 0.96× on Linux, where FoundationXML is libxml2. Parity with libxml2 while building a tree its SAX path never builds is a fine result — but “faster than Foundation’s XML” is not a portable sentence, so it is not said here.

Encoding is a measurement, not a thesis. About 8.28× JSONEncoder at fifty items, 8.02× at two hundred. It was 2.9× until the writers stopped appending to an Array and started owning their buffers.

The decode multiple has an argument behind it. This one does not — it is here so the cost is known and a regression shows up.

The other formats are tree decoders, and the thesis does not transfer

Section titled “The other formats are tree decoders, and the thesis does not transfer”

JSON decodes straight from bytes into your fields. YAML, XML, TOML and property lists reach your struct through RawValue, the format-neutral projection — so there is no Codable boundary being deleted, and no 9× to be had from deleting it. What those numbers measure is a hand-written Swift parser against whatever the ecosystem already offers.

They do not build a node tree on the way any more. Until September 2026 each parser built its own YAML.Node or XML.Element tree, projected that into RawValue, and dropped it; each parser is now generic over what it builds, so the struct door builds RawValue directly and the tree door still builds the tree for callers who ask for one. That is where the end-to-end column below moved from 11.04× to 17.47×.

Format Tree End to end into a struct
YAML 8.28× Yams compose 17.47× Yams YAMLDecoder
XML 2.53× Foundation on macOS, 0.96× on Linux —
TOML 3.69× toml++ 6.92× TOMLKit’s decoder

The pattern in the right-hand column is the familiar one. Where a baseline goes through Codable, the gap widens; where it does not, the gap is parity with C. That is the same finding as the JSON thesis, arrived at from the other direction.

@Schema costs about 84 ms per type at ten fields — roughly 4× Codable. The cost model is 9 ms fixed per type + 7.3 ms per field, so it scales with generated body size rather than with the number of expansions.

That inverts the obvious optimisation: “call the macro less” buys nothing, “emit less code per field” buys everything.

It is a gate rather than a footnote for a simple reason. You replace : Codable with @Schema across a model layer in one commit, and then you wait for a build. That is where you decide whether to keep it — before you have run a single benchmark. CI fails above 100 ms/type.

The opt-ins are what they are for:

Configuration Cost per type, ten fields
@Schema alone ~80 ms
a rule on nearly every field ~116 ms
encodes: true about 5% more
formats: .all ~34 ms more, for the shared YAML/XML/TOML body

Allocation counts, against absolute thresholds: live blocks per decoded value, never wall clock. A timing gate on a shared runner is a flaky test with extra steps. The counter self-checks against closures whose block count is arithmetic, and switches itself off rather than report a number it cannot stand behind.

Instructions, retains, releases and heap bytes are counted exactly too, per call, under Valgrind — on both x86-64 and arm64, every push. Those ratchet: an allocation that appears where there was none fails the build.

Total malloc traffic is measured on Darwin and reported rather than gated. It has no a-priori right answer, and gating it would fail CI the day String changes its growth policy.

Terminal window
cd Benchmarks
swift run -c release CorpusGen # the corpus, deterministic
swift run -c release AssayBench --list # every arm
swift run -c release AssayBench falsification

The source of this table, with how to read each comparison and in whose favour it is arranged, is Benchmarks/RESULTS.md in the repository.