← Back to blog

Deterministic LLM JSON: 5 Normalization Checks Engineers Run in CI

October 5, 2026
Deterministic LLM JSON: 5 Normalization Checks Engineers Run in CI

Run five checks on every piece of LLM-generated JSON before it touches downstream code: constrained decoding at generation time, isolation of the JSON block from surrounding text, syntax repair, safe parsing with a reviver or object hook, and strict schema validation with safety controls. This applies specifically to JSON produced by language models, not to database or document-store normalization. The pipeline section below gives copy-paste examples for each step.


TL;DR:

  • Prioritize schema validation, strict schema enforcement, and input size limits to prevent malformed or malicious JSON from reaching downstream systems.
  • Use a layered pipeline: isolate JSON from extraneous text, repair syntax errors, normalize keys and numeric formats, then validate against a strict schema.
  • Recognize common failure patterns such as truncation, unescaped characters, extra fields, or schema drift, and handle each accordingly with rejection or repair.
  • Log every repair attempt with original input, applied fix, and validation outcome to support incident review and model improvement.
  • Regularly test and monitor validation and repair steps in production, enforce timeouts, and canonicalize JSON before caching to ensure performance at scale.

Datatool
datatool.dev
Repair AI-Generated JSON Reliably
Datatool helps developers repair, validate, and test malformed structured data from language models before it reaches downstream code.
Visit Datatool

Table of Contents

Which normalization technique should you run first?

Start with prevention, then layer repair and validation behind it. No single technique catches everything an LLM produces, so treat this as a pipeline, not a choice.

Structured Outputs or constrained decoding comes first when your model supports it. OpenAI's Structured Outputs guide documents schema-constrained generation that reduces invalid JSON, though models can still produce mistakes or partial refusals under constraint.

Encapsulation and extraction handles the scratchpad problem. Models often wrap JSON in explanation text or markdown fences. Extract the JSON block with a delimiter search or regex before you try to parse anything.

Syntax repair fixes mechanical damage: unbalanced braces, single quotes where double quotes belong, trailing commas, truncated output, and bad escape sequences.

Key and type normalization standardizes naming (snake_case versus camelCase) and numeric precision, since large integers can lose precision when parsed as floating-point numbers.

Validation and safety checks decide what happens next:

  • Reject immediately when required fields are missing and no safe default exists.
  • Attempt deterministic repair when the damage is mechanical and the fix is unambiguous.
  • Never guess at semantic content. A missing brace is fixable. A missing business value is not.

A repair pipeline you can paste into a test suite

The pattern is isolate, repair, parse, validate. Here's a broken response and the fix at each stage.

Say a model returns this:

Sure, here's the result:
{"user_id": 12345678901234567, "name": "Ada Lovelace", "active": true,

That's scratchpad text, a truncated object, and a number too large for safe float handling.

  1. Isolate. Find the first { and the last } in the raw string, or use a ```json fence if your prompt requests one. If no closing brace exists, the extraction step itself signals truncation.
  2. Repair. Count open versus closed braces and brackets. Append what's missing. Strip trailing commas before a closing brace.
  3. Parse with controlled coercion. In JavaScript, JSON.parse accepts a reviver function that runs after parsing:
JSON.parse(repaired, (key, value) => {
  if (key === "user_id" && type of value === "number") {
    return String(value); // preserve precision
  }
  return value;
});

In Python, json.loads takes an object_hook for the same purpose, coercing large integers to strings before they lose precision.

  1. Validate against a strict schema. Use JSON Schema Draft 2020-12 with additionalProperties: false and explicit required fields so drift gets caught, not silently accepted.

The reviver and object_hook only run on JSON that already parses. MDN's documentation is explicit that a reviver cannot prevent a SyntaxError on malformed input, so repair has to happen before parsing, not during it.

Pro Tip: Log every repair with the original input, the repaired output, and a pass or fail flag. When a heuristic fails silently in production, that log is the only way to reconstruct what happened.

Common failure modes and how to fix each one

Most broken LLM JSON falls into a handful of repeatable patterns. Recognizing the pattern tells you whether to repair or reject.

  • Scratchpad text: the model explains itself before or after the JSON. Fix with the isolation step above; never attempt to parse the full response body.
  • Partial or truncated objects: output cuts off mid-field, usually from a token limit. If the cut happened inside a string value, repair is unsafe; retry the request instead of guessing at the missing content.
  • Extra or undocumented fields (schema drift): the model adds fields you didn't ask for. Set additionalProperties: false in your schema so these get rejected or stripped explicitly, rather than silently flowing downstream.
  • Enum and type hallucinations: a field expected to be "pending" or "active" comes back as "Pending Review". Build a normalization map for known variants and reject anything outside it, rather than coercing blindly.
  • Escaping and string corruption: unescaped quotes inside string values break parsing. This is repairable when the corruption is consistent (for example, an unescaped " inside a name field) but should be rejected when the string structure itself is ambiguous.

Our team at datatool.dev has documented these failure patterns in more detail, including which ones are safe to auto-repair and which need a human in the loop.

Validation and schema rules that reduce security risk

Schema validation isn't just about catching malformed data. It's a security boundary. OWASP's LLMSVS v2.0 treats LLM output as untrusted input and calls out schema drift specifically as a path to type confusion vulnerabilities.

  • Set required and additionalProperties: false on every object in your schema. This is the single highest-leverage rule for catching unexpected fields before they reach business logic.
  • Enable format checks (date-time, email, uri) rather than trusting that a field named created_at actually contains a valid timestamp.
  • Serialize large integers as strings when precision matters. JavaScript's number type loses precision above 2^53, and silent truncation is worse than an explicit string field.
  • Never pass raw LLM output into SQL, shell commands, or code generation without parameterization and context-aware encoding, even after it passes JSON validation.
  • Add schema conformance tests to CI using a corpus of known-bad examples, not just happy-path fixtures.

Constrained decoding frameworks currently miss meaningful amounts of complex schema coverage, according to JSONSchemaBench, a benchmark evaluating structured-output frameworks against real-world schema complexity. The benchmark finds tradeoffs between over-constraining generation (which can produce valid but wrong output) and under-constraining it (which lets errors through). That's the core argument for keeping independent validation in your pipeline even when you use Structured Outputs or similar constrained decoding at generation time. Our guide to JSON Schema validation walks through building these rules into a working validator.

Production hardening: limits, logging, and test harnesses

Normalization that works in a notebook can still fail in production without the right guardrails.

  • Enforce a maximum input size before parsing. Unbounded strings sent to JSON.parse or json.loads can tie up CPU on deeply nested or pathological input.
  • Set a parse timeout, especially for recursive or deeply nested structures.
  • Log every repair attempt: the raw input, the repair applied, and whether validation accepted or rejected the result. This is what makes an incident reviewable after the fact.
  • Cap automatic repair retries. After two or three failed attempts, route to manual review instead of looping.
  • Keep a corpus of real broken examples from production as permanent regression tests, not synthetic ones you wrote yourself.

Pro Tip: Store the repair log with a timestamp and the schema version used for validation. Schema versions drift over time, and without that pairing, you can't tell whether a rejected payload was actually bad or just validated against the wrong rules.

Handling nested and deeply recursive JSON structures

Deep nesting breaks normalization logic in ways that flat objects don't. A missing brace three levels deep can cascade into a parser misreading the entire remainder of the document, closing the wrong object and corrupting data that was otherwise valid.

Track bracket and brace depth explicitly during the repair step instead of relying on a single global count. A stack-based approach, where you push on each opening character and pop on each matching close, tells you exactly which level is unbalanced and lets you close only that level rather than guessing at the whole structure.

Stack tracking nested JSON bracket depth

Recursive schemas need the same care. JSON Schema supports $ref for self-referencing definitions, which is how you validate a tree-shaped object (nested comments, organizational charts, file systems) without writing out every possible depth by hand. Set a maximum recursion depth in your validator, separate from the schema itself, so a maliciously or accidentally deep structure can't exhaust memory or stack space during validation.

For repair specifically, work from the innermost unclosed structure outward. Attempting to repair a deeply nested object from the outside in tends to produce a structure that parses successfully but places values at the wrong nesting level, which is a worse failure than a parse error because it fails silently downstream.

When an LLM response involves nested arrays of objects, validate each array element independently against the item schema rather than validating the array as a single blob. This isolates which specific element broke instead of failing the whole response on one bad entry.

Canonical JSON forms and why they matter here

A canonical form means there's exactly one correct way to represent a given piece of data: keys in a fixed sort order, consistent whitespace, consistent number formatting. Canonicalization matters for JSON normalization in two specific ways.

First, it makes repaired output comparable. If you're running regression tests against a corpus of known-bad LLM responses, the expected output has to be in a canonical form or your test will fail on whitespace differences that have nothing to do with correctness. Sort object keys alphabetically and normalize number formatting before diffing two JSON documents in a test harness.

Second, canonical form makes deduplication and caching possible. If your pipeline sees the same logical response represented two different ways (different key order, different number formatting) a cache keyed on the raw string treats them as different entries. Canonicalizing before hashing fixes that.

Canonicalization is not the same as validation. A document can be perfectly canonical and still violate your schema, and a document can satisfy your schema while having inconsistent key ordering. Run canonicalization after repair and before you store or cache a result, and run schema validation independently of both.

Keep the canonical transform deterministic and documented. If two engineers on your team sort keys differently or handle Unicode normalization differently, your "canonical" form isn't canonical at all, and your diffs and caches will quietly disagree with each other.

Tools and libraries built for this problem

General-purpose JSON libraries assume well-formed input. LLM output rarely cooperates, which is why a separate category of tooling exists specifically for repair and normalization of AI-generated JSON.

Projects like llm-structured-output use encapsulation delimiters and JSON acceptors to locate the actual JSON block inside a noisy response, then run schema-aware acceptors against just that block. That matches the isolate-then-validate pattern covered earlier in this article.

For teams that want this handled without maintaining custom heuristics, our JSON repair tool at datatool.dev and the @datatool/json-heal package apply deterministic repair rules to malformed AI output and log every decision for audit. We built this specifically around real-world LLM failures (broken brackets, wrapped responses, partial objects, invalid escaping, truncation, schema drift) rather than synthetic test cases.

Standard library tooling still has a role. Python's json module documentation covers object_hook, allow_nan, and JSONDecodeError behavior that you'll lean on constantly when writing custom coercion logic, and JSON.parse on MDN remains the reference for reviver behavior in JavaScript. Neither library repairs malformed input on its own. Both expect you to hand them something that already parses, which is the gap repair tooling exists to fill.

Our guide to JSON validators in AI pipelines covers where to place validator calls relative to repair steps in a typical service architecture.

Performance considerations when normalizing at scale

Normalization logic that's fast enough for one request per second can become a bottleneck at production volume. A few specific costs tend to dominate.

Regex-based extraction and repair heuristics can run in exponential time on pathological input if the patterns aren't written carefully. Test your extraction regex against deliberately adversarial input (deeply nested fake brackets, repeated near-matches) before trusting it in a hot path.

Schema validation cost scales with schema complexity, not just document size. A schema with many $ref chains or deeply nested oneOf/anyOf branches can make validation itself the slow part of your pipeline, not parsing or repair. Profile validation separately from parsing when you're chasing latency.

Logging every repair attempt, which earlier sections recommend for auditability, adds I/O overhead. Batch log writes asynchronously rather than writing synchronously on the request path, especially if you're processing LLM output at any meaningful volume.

Caching helps most when you canonicalize before hashing, as covered above. Without canonicalization, semantically identical responses with different key ordering or whitespace won't hit the cache, and you'll pay the full repair and validation cost repeatedly for data you've already processed.

Finally, set a hard timeout on the whole normalization pipeline, not just the parse step. A repair heuristic that loops on bracket counting, a validator stuck on a recursive schema, or a regex that backtracks badly can each individually stall a request well past any reasonable latency budget.

Performance considerations when normalizing at scale — overview diagram

Where automation helps and where it still falls short

Constrained decoding and automated repair genuinely cut the volume of broken JSON you have to triage by hand. That's real progress. But JSONSchemaBench shows current frameworks still miss meaningful schema coverage on complex cases, and no repair heuristic can recover information the model never generated.

The practical stance: automate the mechanical fixes, log everything, and keep independent validation in the loop. Testing approaches can treat every repair as something to verify, not trust by default.

— Gregory

datatool.dev: deterministic JSON repair for AI pipelines

We built the JSON repair tool at datatool.dev to drop into the pipeline described above: paste in malformed AI output, get back repaired, validated JSON with a transparent log of exactly what changed. No guessing on ambiguous values, no silent corrections.

Datatool

For teams integrating repair directly into their stack, @datatool/json-heal gives you the same deterministic repair engine as a package, with logs you can wire into your own audit trail or CI checks. Check out Datatool to see how it fits your pipeline.

FAQ

What's the difference between JSON repair and JSON validation?

Repair fixes mechanical syntax problems like missing braces, bad quotes, or trailing commas so a string becomes parseable. Validation checks that already-parsed JSON matches your expected structure, types, and required fields. You need both, run in that order.

Can structured outputs replace schema validation entirely?

No. OpenAI's Structured Outputs reduces invalid JSON at generation time, but JSONSchemaBench shows current frameworks still miss complex schema coverage, and models can still produce mistakes or partial refusals. Independent validation remains necessary as a second layer.

Why does JSON.parse fail on valid-looking LLM output?

Common culprits are trailing commas, single quotes instead of double quotes, unescaped characters inside strings, and truncated output from a token limit. MDN's JSON.parse documentation confirms the reviver callback runs after parsing and cannot fix a syntax error, so repair has to happen before you call parse.

How do I handle large integers that lose precision in JSON?

Serialize large integers as strings in your schema rather than relying on native number parsing, since JavaScript numbers lose precision above 2^53. Use a reviver in JavaScript or an object_hook in Python to coerce these fields during parsing.

Should I reject or repair JSON with extra fields?

Set additionalProperties: false in your schema so extra fields are explicitly rejected rather than silently passed through, which OWASP's LLMSVS flags as a schema drift risk. Reject when the extra field could indicate a type confusion issue; only map or strip it when you've confirmed it's benign.

Sources