Generate taxonomy driven malformed JSON fixtures: truncation, wrapped text, invalid escaping, invisible characters. Run each one through parse, schema validation, deterministic repair, and re-validation. This matters because LLM output breaks parsers in specific, repeatable ways, not random ones. Build fixtures around those patterns, following RFC 8259 syntax rules and OWASP guidance on untrusted output.
TL;DR:
- Malformed JSON fixtures should be deterministic and targeted, simulating specific failure modes such as truncation, invalid escaping, invisible characters, or partial objects.
- Validation pipelines must include parsing, schema validation, length checks, deterministic repair, and re-validation; skipping steps risks undetected bugs or security issues.
- Schema validation alone cannot ensure safety, as dangerous payloads can pass type and structure checks; semantic validation and sanitization are essential.
- Invisible characters, silent truncation, and schema-valid but malicious payloads are common pitfalls that developers often overlook until issues appear in production.
- Using repair tools and logging diffs helps trace failures, improve tests, and ensure robust handling of AI-generated JSON, especially for security-critical applications.
Table of Contents
- What kinds of broken JSON do LLMs actually produce?
- How do you generate realistic malformed JSON fixtures?
- What should your validation pipeline check at each step?
- Why can schema-valid JSON still be dangerous?
- How does deterministic repair fit into your test suite?
- What developers consistently get wrong about LLM JSON
- Try deterministic JSON repair in your own pipeline
- Sources
- FAQ
What kinds of broken JSON do LLMs actually produce?
LLMs don't fail randomly. They fail in a short list of predictable ways, and your test suite should cover each one on purpose.
- Truncation: the model hits
max_tokensmid-array or mid-object, leaving an unclosed bracket or a dangling comma. - Wrapped text: the JSON arrives inside a sentence, a markdown code fence, or an explanation ("Here's your JSON: {...}").
- Invalid escaping: embedded user content breaks quotes or backslashes, especially with code snippets or file paths.
- Partial objects: required fields go missing entirely, a sign of schema drift rather than a syntax error.
- Invisible characters: zero-width spaces, variation selectors, or bad surrogate pairs slip in and pass visually but fail parsers.
- Structural errors: duplicate keys, unexpected
additionalProperties, or a string where the schema expects a number.
Each of these needs its own fixture. A repair tool that handles truncation well might do nothing for invisible characters. RFC 8259 notes that parsers can disagree on duplicate keys and surrogate pairs, which is exactly why you test both, not just one. Mixing failure types into a single fixture hides which code path actually failed.
How do you generate realistic malformed JSON fixtures?
Start from valid JSON and apply deterministic mutations. Random fuzzing finds bugs eventually, but targeted mutation reproduces what LLMs actually do, faster and with less noise in your test output.
- Take a canonical valid response and remove the final closing brace to simulate truncation.
- Insert explanatory text before or after the JSON block to simulate a wrapped response.
- Swap a straight quote inside a string value to break escaping.
- Insert a zero-width character (
) inside a key or value. - Delete a required field to simulate schema drift.
- Truncate at a random byte offset near the end of the string to mimic a token-limit cutoff.
Here's a failure and its fix in Python:
import json
broken = '{"name": "Widget", "price": 19.99, "tags": ["new"'
try:
json.loads(broken)
except json.JSONDecodeError as e:
print(f"Parse failed: {e}")
# Minimal repair: close the array and object
repaired = broken + "]}"
data = json.loads(repaired)
print(data)
The JSONDecodeError tells you exactly where the parser gave up. That position is useful. Log it, don't just catch and discard it.
For invisible characters, build a fixture like this:
payload = '{"id": "abc123", "status": "ok"}'
Then run a sanitizer that strips zero-width characters before validation:
import re
def strip_invisible(text):
return re.sub(r'[-]', '', text)
clean = strip_invisible(payload)
data = json.loads(clean)
Skip the sanitizer step and this fixture parses fine but stores a corrupted key that silently mismatches everywhere else in your system.
Pro Tip: Keep mutated fixtures as golden files in your repo. When a repair function changes, diffing against golden files shows exactly which cases regressed.
Simulate OpenAI Structured Outputs edge cases too: a refusal field instead of content, or an incomplete status with reason max_output_tokens. Our guide to Structured Outputs failure modes covers the specific fields to fixture. Complex nested schemas fail more often than flat ones, a pattern we break down in why complex JSON schemas break structured output.
What should your validation pipeline check at each step?
Run every fixture through the same five-step pipeline. Skipping a step means a whole class of bugs ships unnoticed.
- Parse attempt: try
json.loads()or equivalent; log the exception type and position, not just "failed." - Schema validation: check types, required fields, and
additionalPropertiesagainst your JSON Schema. - Length and completeness checks: flag responses that stop suspiciously close to a token limit.
- Deterministic repair: apply a repair function and record exactly what it changed.
- Re-validation: run the repaired output through the same schema check before it touches your database or downstream code.
Schema-valid doesn't mean safe. A field can pass every type check and still carry an injection string or a path traversal attempt. Add semantic checks, like allowed value lists or context-aware escaping, on top of schema validation. Our practical guide to validating AI generated data and JSON Schema validation walkthrough go deeper on both layers.
In CI, assert on three things: the exception type the parser raised, the diff the repair step produced, and final schema conformance. Fail the build on any unrepaired parse error, but route ambiguous repairs to a human review queue instead of silently accepting them.
Why can schema-valid JSON still be dangerous?
Passing schema validation only confirms structure and types. It says nothing about intent or content, and that gap is where security problems live.
- Treat every LLM output as untrusted input, the same way you'd treat a form submission from an anonymous user, per OWASP's guidance for LLM applications.
- Include injection-like strings in fixtures: SQL fragments, shell metacharacters, script tags, to confirm your sanitizers actually catch them.
- Use parameterized queries and context-aware encoding before any output touches a database or a browser.
- Strip zero-width and variation-selector characters at ingest, before validation runs, not after.
OWASP's LLM application guidance explicitly lists zero-width and variation-selector characters as a risk category, recommending they get stripped at the ingest boundary rather than caught downstream, per OWASP's Top 10 for LLM Applications. A test suite that never includes these characters will never catch a sanitizer that misses them.
How does deterministic repair fit into your test suite?
There are tools for repairing, validating, and testing AI-generated structured data, including the failure types above: truncation, wrapped responses, invalid escaping, and schema drift. A minimal test using a JSON repair library looks like a repair call followed by a schema check, with the repair diff logged for traceability. Capturing that diff alongside the final validation result makes a repair auditable instead of a black box.

What developers consistently get wrong about LLM JSON

Three things trip up almost every team I've looked at. First, invisible characters. They don't show up in a terminal or a diff tool, so they get missed until a key lookup fails in production. Second, silent truncation. Third, schema-valid but dangerous values: a string field that matches its type but carries a payload your schema was never designed to catch.
Keep your fixture set narrow. Ten well-chosen cases covering real failure modes beat a hundred random ones. For deeper debugging patterns, our post on detecting AI output errors walks through specific detection strategies, and defending against schema-valid injection covers the security side in more depth.
— Gregory
Try deterministic JSON repair in your own pipeline
Paste messy JSON. Get valid JSON back, with a log of exactly what changed. That's the core of the JSON repair tool, built for the truncation, wrapping, and escaping failures covered above.
The fastest way to start: drop a failing fixture into the tool, confirm the repair, then wire @datatool/json-heal into one unit test. Once that test passes reliably, expand to your full fixture set. See how the repair engine works or check the full product overview before you commit a plan.
Sources
FAQ
What is JSON test data generation for LLM outputs?
It means building malformed, truncated, or partial JSON fixtures that mimic real failures from language models, then using them to test repair and validation pipelines. This differs from generating synthetic valid data for general app testing: the goal here is deliberately broken input, not clean mock records.
How do I simulate OpenAI Structured Outputs failures?
Simulate both a refusal field in place of content and an incomplete status with reason max_output_tokens, since OpenAI's Structured Outputs documentation confirms both states can occur even with constrained decoding. Your test fixtures should cover each state separately, since downstream code needs to handle them differently.
Why does schema validation alone not catch dangerous payloads?
Schema validation checks types and required fields, not content intent, so a string field can pass validation while carrying an injection attempt. OWASP's guidance for LLM applications recommends treating all output as untrusted and adding semantic checks beyond schema conformance.
What should I log when testing JSON repair functions?
Log the original parse exception type and position, the exact diff the repair function produced, and the final schema validation result. This gives you an audit trail for debugging and lets you catch regressions when a repair function changes behavior.
Can I rely on a second LLM call to validate the first one's output?
No. OWASP's guidance for LLM applications specifically advises against relying on a secondary LLM call for validation, recommending trusted application code with schema checks and length checks instead, per the OWASP Top 10 for LLM Applications.

