LLM Structured Outputs: Beyond Valid JSON
A model can return valid JSON and still cause an application to store incorrect data. Parsing checks whether software can read the response. Field constraints check its shape. Business rules and source evidence determine whether the content is usable. Structured LLM output belongs inside this validation pipeline.
Contents
Prompts and output constraints serve different purposes
A prompt explains the task and field meanings, such as deciding whether supplied material contains an answer. Output constraints restrict the allowed form. The official vLLM guide provides structured generation through JSON Schema, fixed choices, regular expressions and grammars, suited to different tasks.
Writing “return JSON” in a prompt is different from enabling a server-side output constraint. Conversely, a structural constraint does not know business facts. A correctly shaped value can still be a misreading or a guess.
Design a small response contract
Suppose a retrieval service permits exactly two fields: status and evidence. Status is found or not_found; evidence is an array of source-ID strings. Add a business rule: found requires at least one ID, while not_found requires an empty array.
| Response | Issue |
|---|---|
| found + empty array | Types may be correct, but the business relationship contradicts itself |
| found + [doc-1] | After structural checks, verify that doc-1 exists and supports the conclusion |
| not_found + empty array | Missing evidence does not prove that no answer exists |
An explicit missing-evidence outcome prevents incomplete information from being hidden inside an apparently complete success object. Source IDs should not be invented by the model. Restrict candidate sources and check them again in the application.
Keep checking after parsing
This standard-library example implements the small contract. Seven cases were executed, covering valid responses, a bad enum, an incorrect type, an extra field, a business contradiction and broken JSON. It is not a general JSON Schema validator and performs neither model inference nor a vLLM run.
import json
def validate(obj):
if not isinstance(obj, dict) or set(obj) != {'status', 'evidence'}:
raise ValueError('expected exactly status and evidence')
if obj['status'] not in ['found', 'not_found']:
raise ValueError('unknown status')
if not isinstance(obj['evidence'], list) or not all(type(x) is str for x in obj['evidence']):
raise ValueError('evidence must be a list of strings')
if (obj['status'] == 'found') != bool(obj['evidence']):
raise ValueError('status and evidence disagree')
return obj
for text, expected in [
('{"status":"found","evidence":["doc-1"]}', True),
('{"status":"not_found","evidence":[]}', True),
('{"status":"found","evidence":[]}', False),
('{"status":"maybe","evidence":[]}', False),
('{"status":"found","evidence":"doc-1"}', False),
('{"status":"not_found","evidence":[],"extra":1}', False),
('{broken', False),
]:
try:
validate(json.loads(text)); actual = True
except (ValueError, TypeError):
actual = False
assert actual == expected
print('7 validation cases passed')
It prints “7 validation cases passed”. The final relationship check does not inspect doc-1’s content; that requires the actual source. Cross-Encoder reranking can help order retrieval candidates, but its scores do not establish factual truth either.
Put business rules where they can be tested
Some field relationships can be encoded in Schema constraints or kept in a separate business function. The choice depends on supported Schema features and whether a rule needs a database or outside evidence. Either way, maintain executable failure cases instead of relying only on prose.
Define extra-field handling, defaults for missing fields and the meaning of empty arrays explicitly. Quietly changing maybe to found, or replacing a parse failure with a default success, erases genuine failure information. Bound retries and distinguish formatting failures, semantic failures and insufficient evidence.
Check the inference integration
Verify the installed version, request parameters and constraint support of the selected backend. The vLLM guide includes parameter migration notes, so copied older fields should not be assumed compatible. This article supplies no untested inference deployment command and does not claim every Schema works on every backend.
Check response completeness as well. A partial object in a stream is not a final result. Timeouts or length-limit truncation should enter a failure path. Use explicit denominators when reporting parsing, business-rule and evidence-verification pass rates; one undifferentiated success rate can conceal failures.
Structured output helps connect a model to software. Trust still depends on constraints, validation and evidence working together. A parseable object is the beginning of the process.
References
Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the originally planned October 9, 2026, 14:00 article date. This original contract-validation example is not a model benchmark.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


