LLM Structured Outputs: Beyond Valid JSON

黎 浩然/ 9 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

A model can return valid JSON and still cause an application to store incorrect data. Parsing checks whether software can read the response. Field constraints check its shape. Business rules and source evidence determine whether the content is usable. Structured LLM output belongs inside this validation pipeline.

Contents
  1. Prompts and output constraints serve different purposes
  2. Design a small response contract
  3. Keep checking after parsing
  4. Put business rules where they can be tested
  5. Check the inference integration
  6. References

中文版

Prompts and output constraints serve different purposes

A prompt explains the task and field meanings, such as deciding whether supplied material contains an answer. Output constraints restrict the allowed form. The official vLLM guide provides structured generation through JSON Schema, fixed choices, regular expressions and grammars, suited to different tasks.

Writing “return JSON” in a prompt is different from enabling a server-side output constraint. Conversely, a structural constraint does not know business facts. A correctly shaped value can still be a misreading or a guess.

Parse JSON
Check fields
Check business rules
Verify evidence
Original validation pipeline. Passing a stage does not imply that later stages pass.

Design a small response contract

Suppose a retrieval service permits exactly two fields: status and evidence. Status is found or not_found; evidence is an array of source-ID strings. Add a business rule: found requires at least one ID, while not_found requires an empty array.

Response Issue
found + empty array Types may be correct, but the business relationship contradicts itself
found + [doc-1] After structural checks, verify that doc-1 exists and supports the conclusion
not_found + empty array Missing evidence does not prove that no answer exists

An explicit missing-evidence outcome prevents incomplete information from being hidden inside an apparently complete success object. Source IDs should not be invented by the model. Restrict candidate sources and check them again in the application.

Keep checking after parsing

This standard-library example implements the small contract. Seven cases were executed, covering valid responses, a bad enum, an incorrect type, an extra field, a business contradiction and broken JSON. It is not a general JSON Schema validator and performs neither model inference nor a vLLM run.

import json

def validate(obj):
    if not isinstance(obj, dict) or set(obj) != {'status', 'evidence'}:
        raise ValueError('expected exactly status and evidence')
    if obj['status'] not in ['found', 'not_found']:
        raise ValueError('unknown status')
    if not isinstance(obj['evidence'], list) or not all(type(x) is str for x in obj['evidence']):
        raise ValueError('evidence must be a list of strings')
    if (obj['status'] == 'found') != bool(obj['evidence']):
        raise ValueError('status and evidence disagree')
    return obj

for text, expected in [
    ('{"status":"found","evidence":["doc-1"]}', True),
    ('{"status":"not_found","evidence":[]}', True),
    ('{"status":"found","evidence":[]}', False),
    ('{"status":"maybe","evidence":[]}', False),
    ('{"status":"found","evidence":"doc-1"}', False),
    ('{"status":"not_found","evidence":[],"extra":1}', False),
    ('{broken', False),
]:
    try:
        validate(json.loads(text)); actual = True
    except (ValueError, TypeError):
        actual = False
    assert actual == expected
print('7 validation cases passed')

It prints “7 validation cases passed”. The final relationship check does not inspect doc-1’s content; that requires the actual source. Cross-Encoder reranking can help order retrieval candidates, but its scores do not establish factual truth either.

Put business rules where they can be tested

Some field relationships can be encoded in Schema constraints or kept in a separate business function. The choice depends on supported Schema features and whether a rule needs a database or outside evidence. Either way, maintain executable failure cases instead of relying only on prose.

Define extra-field handling, defaults for missing fields and the meaning of empty arrays explicitly. Quietly changing maybe to found, or replacing a parse failure with a default success, erases genuine failure information. Bound retries and distinguish formatting failures, semantic failures and insufficient evidence.

Check the inference integration

Verify the installed version, request parameters and constraint support of the selected backend. The vLLM guide includes parameter migration notes, so copied older fields should not be assumed compatible. This article supplies no untested inference deployment command and does not claim every Schema works on every backend.

Check response completeness as well. A partial object in a stream is not a final result. Timeouts or length-limit truncation should enter a failure path. Use explicit denominators when reporting parsing, business-rule and evidence-verification pass rates; one undifferentiated success rate can conceal failures.

Structured output helps connect a model to software. Trust still depends on constraints, validation and evidence working together. A parseable object is the beginning of the process.

References

vLLM: Structured Outputs

Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the originally planned October 9, 2026, 14:00 article date. This original contract-validation example is not a model benchmark.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*