Reference
Performance
Every row below came off one arm64 Mac: warm, -O, minimum of five rounds, on the corpus in
the repository. None of them is a claim about your machine, or about another platform. Linux
and x86-64 get their own sections in Benchmarks/RESULTS.md, and where a number flips across
platforms this page says so rather than quoting the flattering half.
| Arm | Number | Against |
|---|---|---|
| struct decode, full corpus | 8.75× mean over 25 files (4.95–19.83) | JSONDecoder |
| prefix decode + unknown-key skip | 5.52× over 45 files (2.65–8.22) | JSONDecoder |
| generic value model | 3.31× over 75 files | JSONSerialization |
| falsification arm (apimodel, 5 sizes) | 5.21× mean (7.94× float-dense) | JSONDecoder |
| vs ZippyJSON (simdjson + Codable) | 3.20–3.66× faster, over four sizes | ZippyJSON, which is 1.62–1.82× over Foundation here |
| vs yyjson, use-case shape | 0.71× (loses) | yyjson parse + extraction |
| vs yyjson, float-dense | 0.71× (loses) | same |
| vs yyjson, DOM vs DOM | 0.16× (loses) | yyjson_read |
| YAML node parse | 8.28× | Yams compose |
| YAML struct decode | 17.47× | Yams YAMLDecoder |
| XML tree parse | 2.53× (macOS; 0.96× on Linux, 2026-08-18) | Foundation XMLParser |
| TOML node parse | 3.69× | toml++ via TOMLKit |
| TOML struct decode | 6.92× | TOMLKit TOMLDecoder |
| Date fields | 5.58× mean over 5 sizes | JSONDecoder + .iso8601 |
| binary plist | 4.33× | Foundation PropertyListDecoder |
| XML plist | 1.29× | Foundation PropertyListDecoder |
| union vs its variant | 1.19× (the tag scan) | the variant decoded directly |
| @Inline vs nesting | 0.85× (faster) | the nested @Schema it replaces |
| @Wraps vs @Validate | 1.98× (slower) | the plain field + rule it is sugar for |
| encoding, 50 / 200 items | 8.28× / 8.02× | JSONEncoder |
| cold start, 60 types | 6.1× first decode (median); 5.2× steady | JSONDecoder |
| multi-megabyte documents | 8.41–9.02×, ~1,020 MB/s, flat | JSONDecoder |
| total allocations, 50 items | 159 against Foundation’s 377 | JSONDecoder |
| T.validate(_:) | 40 ns per value, 1 block; 47 ns/row batched, 0.11× a decode | — |
| live allocations, apimodel-8k struct | gated, PASS | absolute thresholds |
| compile time, 10 fields | 84.4 ms/type (budget 100) | Codable: 4.03× |
What the big number means
Section titled “What the big number means”Assay does not need to beat simdjson. It needs to not have a KeyedDecodingContainer.
That is the whole thesis. ZippyJSON bolted simdjson — the fastest JSON parser in existence —
onto Decodable and got 1.38× over Foundation. Apple’s own prototype, which changes nothing
about parsing and only deletes the container protocol, reports about 6×. Roughly 83% of a
Swift decode is the Codable boundary, and a macro deletes it at compile time.
Which is why Assay comes out 3.2–3.7× faster than ZippyJSON here. Scalar Swift, no SIMD anywhere, against simdjson underneath. The container is the only structural difference.
And ZippyJSON is not being hobbled: it measures 1.62–1.82× over Foundation on this machine, better than the 1.38× it was originally cited at.
What it does not mean
Section titled “What it does not mean”It is not faster than C. Against yyjson, Assay loses: 0.71× on the use-case shape, 0.16× building a tree. Those rows are in the table above, published rather than omitted, because a benchmark page that lists only its wins is an advertisement.
Those two rows measure different things, and the gap between them is the interesting part.
The use-case arm makes yyjson do the whole job: parse, then walk the tree and pull the fields out. A document you still have to walk is not a decoded value. There Assay lands at about two thirds of hand-tuned C, which is a fine place to be.
The tree-against-tree arm is where it loses about 6×, and it compares two different products.
yyjson writes one arena with its strings pointing into your input buffer. JSON.Value is a
Swift enum tree of real Strings and Arrays you can store, compare and hash. Closing that
gap means giving up all three.
The value-model path has no thesis behind it. JSON.Value measures about 3.3× over
JSONSerialization — comfortably ahead of the thing you would otherwise reach for, and
nowhere near the struct path, because building a tree has no Codable boundary to delete. Use
@Schema when you know the shape; that is where the argument applies.
XML is a Darwin-only claim. Assay’s XML parser measures 2.53× over Foundation on macOS
and 0.96× on Linux, where FoundationXML is libxml2. Parity with libxml2 while building a
tree its SAX path never builds is a fine result — but “faster than Foundation’s XML” is not
a portable sentence, so it is not said here.
Encoding is a measurement, not a thesis. About 8.28× JSONEncoder at fifty items, 8.02×
at two hundred. It was 2.9× until the writers stopped appending to an Array and started
owning their buffers.
The decode multiple has an argument behind it. This one does not — it is here so the cost is known and a regression shows up.
The other formats are tree decoders, and the thesis does not transfer
Section titled “The other formats are tree decoders, and the thesis does not transfer”JSON decodes straight from bytes into your fields. YAML, XML, TOML and property lists reach
your struct through RawValue, the format-neutral projection — so there is no Codable
boundary being deleted, and no 9× to be had from deleting it. What those numbers measure is a
hand-written Swift parser against whatever the ecosystem already offers.
They do not build a node tree on the way any more. Until September 2026 each parser built
its own YAML.Node or XML.Element tree, projected that into RawValue, and dropped it;
each parser is now generic over what it builds, so the struct door builds RawValue directly
and the tree door still builds the tree for callers who ask for one. That is where the
end-to-end column below moved from 11.04× to 17.47×.
| Format | Tree | End to end into a struct |
|---|---|---|
| YAML | 8.28× Yams compose |
17.47× Yams YAMLDecoder |
| XML | 2.53× Foundation on macOS, 0.96× on Linux | — |
| TOML | 3.69× toml++ | 6.92× TOMLKit’s decoder |
The pattern in the right-hand column is the familiar one. Where a baseline goes through
Codable, the gap widens; where it does not, the gap is parity with C. That is the same
finding as the JSON thesis, arrived at from the other direction.
Compile time is the second axis
Section titled “Compile time is the second axis”@Schema costs about 84 ms per type at ten fields — roughly 4× Codable. The cost
model is 9 ms fixed per type + 7.3 ms per field, so it scales with generated body size
rather than with the number of expansions.
That inverts the obvious optimisation: “call the macro less” buys nothing, “emit less code per field” buys everything.
It is a gate rather than a footnote for a simple reason. You replace : Codable with
@Schema across a model layer in one commit, and then you wait for a build. That is where
you decide whether to keep it — before you have run a single benchmark. CI fails above
100 ms/type.
The opt-ins are what they are for:
| Configuration | Cost per type, ten fields |
|---|---|
@Schema alone |
~80 ms |
| a rule on nearly every field | ~116 ms |
encodes: true |
about 5% more |
formats: .all |
~34 ms more, for the shared YAML/XML/TOML body |
What is gated in CI
Section titled “What is gated in CI”Allocation counts, against absolute thresholds: live blocks per decoded value, never wall clock. A timing gate on a shared runner is a flaky test with extra steps. The counter self-checks against closures whose block count is arithmetic, and switches itself off rather than report a number it cannot stand behind.
Instructions, retains, releases and heap bytes are counted exactly too, per call, under Valgrind — on both x86-64 and arm64, every push. Those ratchet: an allocation that appears where there was none fails the build.
Total malloc traffic is measured on Darwin and reported rather than gated. It has no
a-priori right answer, and gating it would fail CI the day String changes its growth
policy.
Reproducing any of it
Section titled “Reproducing any of it”cd Benchmarksswift run -c release CorpusGen # the corpus, deterministicswift run -c release AssayBench --list # every armswift run -c release AssayBench falsificationThe source of this table, with how to read each comparison and in whose favour it is
arranged, is Benchmarks/RESULTS.md in the repository.