Assay
JSON
JSON is the one format Assay does not build a tree for. The macro writes decode code for your type at compile time, and that code reads the bytes and fills your fields. Nothing in between.
That is where the speed comes from, and it is the only format where the speed argument applies at all. Everything else on this site parses to a value model first.
@Schema(keys: .snakeCase)struct Article { var title: String var link: String var readingMinutes: Int var tags: [String] = []}No formats: argument needed. JSON is the default, and the only one you get for free.
{"title": "On carets", "link": "https://example.com/carets", "reading_minutes": 4, "tags": ["errors"]}Article(title: "On carets", link: "https://example.com/carets", readingMinutes: 4, tags: ["errors"])Nesting, arrays, dictionaries
Section titled “Nesting, arrays, dictionaries”A nested @Schema type is just a field. So is an array of them, and so is a dictionary
with String keys.
@Schema(keys: .snakeCase) struct Server { var host: String; var port: Int; var tls: Bool = true }
@Schema(keys: .snakeCase)struct Cluster { var name: String var servers: [Server] var labels: [String: String] = [:]}{ "name": "eu-prod", "servers": [ {"host": "a.internal", "port": 8080}, {"host": "b.internal", "port": 8081, "tls": false} ], "labels": {"tier": "prod", "team": "platform"}}Cluster(name: "eu-prod", servers: [Server(host: "a.internal", port: 8080, tls: true), Server(host: "b.internal", port: 8081, tls: false)], labels: ["team": "platform", "tier": "prod"])Note tls on the first server. It is absent in the document and true in the result,
because the declaration said so. See presence for the five states and
what each one means.
Numbers are checked, not coerced
Section titled “Numbers are checked, not coerced”JSON distinguishes 1 from 1.5 from "1", so Assay does too. A Double field accepts
an integer literal, because every JSON integer is a valid number. Nothing else widens.
@Schema struct Metrics { var count: Int32; var ratio: Double; var enabled: Bool }{"count": 1.5, "ratio": "half", "enabled": "yes"}metrics.json:1:11: error: count must be an integer, found 1.5 1 │ {"count": 1.5, "ratio": "half", "enabled": "yes"} │ ^
metrics.json:1:25: error: ratio must be a number, found "half" 1 │ {"count": 1.5, "ratio": "half", "enabled": "yes"} │ ^
metrics.json:1:44: error: enabled must be a boolean, found "yes" 1 │ {"count": 1.5, "ratio": "half", "enabled": "yes"} │ ^
3 errorsThree mistakes, three carets, one pass. The third is the interesting one: "yes" is a
string, and a string is not a boolean, so it is an error rather than a quiet true.
Overflow is checked against the declared width, not against Int64:
{"count": 99999999999, "ratio": 0.5, "enabled": true}metrics.json:1:22: error: count must be an integer 1 │ {"count": 99999999999, "ratio": 0.5, "enabled": true} │ ^
1 errorcount is an Int32. The value fits in an Int64 comfortably and the document is
perfectly well-formed, so this is a schema error rather than a parse error.
If you want "8080" to decode into an Int, that is
coerceScalars and it is opt-in per type or per field.
Duplicate keys
Section titled “Duplicate keys”RFC 8259 leaves duplicates undefined. Assay takes the last one:
{"title": "first", "title": "second", "link": "l", "reading_minutes": 1}Article(title: "second", link: "l", readingMinutes: 1, tags: [])The value model is the other answer. JSON.Value keeps every member in document order,
duplicates included, because throwing one away silently is worse than handing you both.
When the document itself is broken
Section titled “When the document itself is broken”Parse errors come from the parser and read differently from schema errors. They still carry a caret.
A trailing comma, which is the one everybody hits:
{"title": "x", "link": "y", "reading_minutes": 1,}bad.json:1:50: error: is not well-formed: expected a key in double quotes 1 │ {"title": "x", "link": "y", "reading_minutes": 1,} │ ^
1 errorAn unquoted key, which is JavaScript and not JSON:
{title: "x"}bad.json:1:2: error: is not well-formed: expected a key in double quotes 1 │ {title: "x"} │ ^
1 errorA truncated document, where there is no byte to point at because the bytes ran out:
{"title": "x", "link":bad.json:1:22: error: is not well-formed: the input ended where ',' or '}' was expected 1 │ {"title": "x", "link": │ ^
bad.json: error: link must be a string
2 errorsThat last render is worth reading twice. The schema reports what it was missing and the parser reports that the document never ended, because both are true and you would want to know both.
Unknown keys
Section titled “Unknown keys”By default they are ignored, which is what you want for an API that adds fields without telling you. Three other policies exist:
@Schema(unknownKeys: .warn) // decode, but say so@Schema(unknownKeys: .reject) // an error, with a did-you-mean@Schema(unknownKeys: .collect) // into an @Extras dictionary.reject and .warn both run a Damerau edit-distance check against the keys the schema
knows, so a typo gets named rather than merely counted. Keys has the
whole story, including aliases and paths.
What you can hand it
Section titled “What you can hand it”[UInt8] is the real overload and String is a convenience that copies into one. Data
lives in AssayFoundation and is decoded where it already is:
import AssayFoundation
try Report.parse(json: bytes) // [UInt8]try Report.parse(json: text) // String — copied to UTF-8 for youtry Report.parse(json: data) // Data — no copy of the inputArray(data) is what you would otherwise write, and it copies the whole document before the
parse even starts. That copy then sits there for the length of the parse. Skipping it saves
one allocation per decode whatever the size, plus 1–3.5% of the time on documents from 0.2 to
8.3 MB — the memory is the real win.
One wrinkle: after a clean decode from Data, d.source is empty. A Data’s bytes are
only valid for the length of the call, so Assay copies them only when an issue or warning
needs a caret. You handed the Data over, so you still have it — and failures render exactly
as they do from an array.
Large documents
Section titled “Large documents”import AssayFoundationlet report = try Report.parse(mmapped: url)Maps the file and decodes in place rather than reading it into Data first — the kernel
pages it in as the parse walks it. On a large document that is a fraction of the memory
footprint and meaningfully faster, and errors still render carets straight out of the
mapping. Throughput stays flat into the multi-megabyte range.
When you do not know the shape
Section titled “When you do not know the shape”JSON.Value is the hand-walkable model. Ordered members, duplicates preserved, integers
and doubles kept distinct.
{"kind": "batch", "items": [{"id": 1}, {"id": 2}], "meta": null}v["kind"]?.string → batchv["items"]?[1]?["id"]?.int → 2v["meta"] → nullv["items"]?.array?.count → 2v["absent"] → nilSubscripts are optional-chaining all the way down, so a wrong guess at any level gives you
nil rather than a trap.
One honest note, because it is the opposite of the rest of this page: the value model is
not the fast path. Building a tree has no Codable boundary to delete, so the argument
that makes @Schema fast does not apply. It measures about 3.3× JSONSerialization and
loses badly to a C DOM parser. Performance has the numbers and
the reasoning. When you know the shape, declare it.