When vLLM structured output stops working, the symptom can be HTTP 200 with unconstrained text rather than an error. The guided_json family was removed in v0.12.0, but vLLM's request models accept unknown fields by design, so a legacy request still returns 200 while the schema is dropped; v0.30.0 (September 22, 2026) adds a warning, and it is logged once per process for each distinct set of removed fields. With a --reasoning-parser configured, the grammar waits for the end of reasoning, and on /v1/completions a raw prompt never ends it: in #57726 a Qwen3-0.6B request that returned valid JSON without the parser returned an unclosed object with finish_reason length once qwen3 was enabled. The default auto backend falls back from xgrammar only when its feature check flags the schema or xgrammar cannot compile it, and multi-branch allOf (#56556) and type-less string constraints (#57550) are not on that list, both open as of September 30, 2026. Speculative decoding adds a louder set: Failed to advance FSM 500s and, in one report, 249,698 trailing spaces behind an HTTP 200. Start by grepping your client code for guided_ and running a four-leg canary with an enum value the model would never write unprompted.
vLLM structured output not working does not have to look like an error. It can look like HTTP 200. In issue #53975, filed on August 27, 2026 against a 0.26.1rc1 nightly serving Gemma 4 31B on two A100s, a request carrying guided_json came back with status 200, 12 completion tokens and the content alpha=1, beta=2, gamma=3. Not JSON. The same schema sent through response_format on the same process returned {"alpha": 1, "beta": 2, "gamma": 3}. A second reporter hit the identical behaviour on 0.28.0 stable.
That is a design decision, not a crash. vLLM's request models accept unknown fields because the OpenAI API does, and guided_json has been an unknown field since v0.12.0. The same shape repeats across the stack: a reasoning parser that never opens the grammar, a schema keyword the backend skips, a speculative decoder that desynchronises the grammar mid-stream. Some of these fail loudly with a 500. The dangerous ones return a normal-looking 200.
This is the troubleshooting map for v0.30.0, released September 22, 2026. How constrained decoding masks tokens is covered in our primer on constrained decoding and streaming JSON, and this post assumes it. The evidence here is the v0.30.0 source and docs, the release notes, and GitHub issue reports. Every number from an issue belongs to its reporter and carries the version they ran.
Why vLLM returns HTTP 200 without enforcing your schema
The base class every OpenAI-compatible request inherits from in v0.30.0 is declared with model_config = ConfigDict(extra="allow"), with a source comment noting that the OpenAI API allows extra fields. Any key the server does not recognise is logged at DEBUG as "The following fields were present in the request but ignored" and dropped. There is no 400 path for unknown fields. A misspelled key, a field from an older release, or a field from another engine's API all produce the same outcome: the request runs without it.
For structured output this inverts the usual contract. A client that sends a schema and receives a 200 has learned only that the server generated some tokens. It has not learned that a grammar constrained them. The HTTP status, the finish reason and the token count all look the same whether the schema was enforced or ignored, which makes this the serving-layer version of the problem in our post on detecting false success in AI agents: the component reporting success is not the component that knows.
The rule that follows is short. Parse and validate every structured response after decoding, against the same schema you sent, and treat a validation failure as an incident rather than a retry. Constrained decoding reduces how often that check fires. It does not remove the need for it.
vLLM structured output errors: symptom, cause and fix by version
Each row carries the version the reporter ran and the issue state as of September 30, 2026. The speculative-decoding rows all predate v0.30.0, and whether they still reproduce there has not been established. v0.30.0 does include PR #51450, which strips speculative-decoding padding tokens before they reach the grammar; despite its title, it does not address the CPU crash or the FSM rejections below.
Reporters' numbers make the speculative-decoding rows concrete. In #52620 a Qwen3-0.6B server on one L40S returned 3 successes and 21 HTTP 500s for 24 requests at concurrency 8, and 24 of 24 successes with speculative decoding off. In #49210 the engine stayed unresponsive for 8.5 hours while the service reported healthy. #51660, on 0.26.0, describes a Kimi-K2.6 plus EAGLE3 setup where no backend was viable: xgrammar desynchronised and killed the engine, outlines leaked threads until OOM, and guidance rejected the slow tokenizer. Our post on why speculative decoding can run slower than baseline covers tuning; for structured traffic, the simplest control is a replica that serves it with speculative decoding off.
The Failed to advance FSM 500s are at least visible. In #49694, on 0.25.1 with ngram_gpu drafting and async scheduling, 178 of 200 chat requests failed with HTTP 500 at concurrency 4, while /v1/completions returned zero errors and mean response length fell from 264 characters at concurrency 1 to 24 at concurrency 12 to 16. Same fault, one endpoint loud, the other silently truncating.
| Symptom | Cause | Reported on | State | Fix or workaround |
|---|---|---|---|---|
200, free text, request uses guided_json | Field removed in v0.12.0, accepted as an unknown key | 0.26.1rc1 and 0.28.0 (#53975) | Closed Sept 4 by a warning-only change (#54285) | Move to response_format or structured_outputs |
200, free text on /v1/completions with --reasoning-parser | Grammar waits for reasoning to end; a raw prompt never ends it | 0.27.1 (#57726); same gate in v0.30.0 source | Open | Use /v1/chat/completions |
200, free text from kimi_k3 when it skips its think channel | Parser never reports reasoning end without a think marker | main, September 2026 (#57714) | Open; fix PR #57718 open | enable_in_reasoning, or wait for #57718 |
200, JSON violating a multi-branch allOf | Not in xgrammar's unsupported-feature check, so no fallback | 0.28.1rc0 (#56556) | Open; fix PR #56557 open | Flatten allOf, or pin guidance |
200, JSON violating pattern plus maxLength on a property with no type | Feature check keys off an explicit type | main, September 2026 (#57550) | Open | Add "type": "string" |
| 200, truncated array then 249,698 whitespace characters | XGrammar desync under DSpark speculative decoding with streaming json_schema | 0.27.1 (#55313) | Open | No speculative decoding on structured traffic |
500, Failed to advance FSM | Grammar rejects a token the engine tries to commit: ngram_gpu drafting in #49694 and #52620; in #53181 the original tool-calling request still completed, and a commenter's 500s on 0.25.1 with speculative decoding off stopped under Model Runner V1 | 0.25.1 (#49694), 0.27.1 (#52620), 0.27.2rc1 nightly (#53181) | Open | No speculative decoding on structured traffic; for the #53181 commenter, VLLM_USE_V2_MODEL_RUNNER=0 |
| 500 with an empty error message | Malformed schema surfaces as InternalServerError | 0.27.1 (#57725) | Open | Validate schemas before sending |
| Engine crash on the CPU backend, all later requests fail | pin_memory=True without an accelerator | 0.25.1 (#51283) | Open | Avoid structured outputs on CPU builds |
| Engine livelock at 100% CPU, service still reports running | MTP speculative decoding with xgrammar, regression from v0.24.0 | 0.25.1 (#49210) | Open | Liveness probe that sends a real structured request |
Migrating off guided_json: removed in v0.12.0, still accepted
PR #29326, merged November 25, 2025, completed the scheduled removal of the guided_* fields, and the v0.30.0 docs state they "were removed in v0.12.0". The docs give a one-to-one mapping.
This request body returns 200 on every release since v0.12.0, and its output is not constrained:
{
"model": "extractor",
"messages": [{"role": "user", "content": "Extract the invoice fields."}],
"guided_json": {"type": "object", "properties": {"total": {"type": "number"}}, "required": ["total"]}
}Either of these two bodies enforces the schema. The first is the OpenAI shape, the second is vLLM-native:
{
"model": "extractor",
"messages": [{"role": "user", "content": "Extract the invoice fields."}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "invoice",
"schema": {"type": "object", "properties": {"total": {"type": "number"}}, "required": ["total"]}
}
}
}{
"model": "extractor",
"messages": [{"role": "user", "content": "Extract the invoice fields."}],
"structured_outputs": {"json": {"type": "object", "properties": {"total": {"type": "number"}}, "required": ["total"]}}
}v0.30.0 added one piece of visibility through PR #54285, merged September 4, 2026. When a request contains any of the six removed fields, the server logs this message (the list is the sorted set of fields it found):
Request contains the removed guided-decoding field(s) ['guided_json'], which are ignored; output will NOT be constrained. Use `structured_outputs` (or `response_format`) instead; see docs/features/structured_outputs.md.
Three properties of that warning matter. It is not in v0.29.0 or earlier. It is emitted with warning_once, which deduplicates on the message and its arguments, so each distinct set of removed fields appears one time per server process and later requests with the same set only reach DEBUG. And the PR states there is no behavioural change: the request still returns 200 with unconstrained output. A flag will not fix this; only a client change will. Find every caller before the upgrade, not after:
grep -rnE 'guided_(json|regex|choice|grammar|decoding_backend|whitespace_pattern)' .
| Removed field | Replacement in the request body | Offline equivalent |
|---|---|---|
guided_json | "structured_outputs": {"json": ...} | StructuredOutputsParams(json=...) |
guided_regex | "structured_outputs": {"regex": ...} | StructuredOutputsParams(regex=...) |
guided_choice | "structured_outputs": {"choice": ...} | StructuredOutputsParams(choice=...) |
guided_grammar | "structured_outputs": {"grammar": ...} | StructuredOutputsParams(grammar=...) |
guided_whitespace_pattern | "structured_outputs": {"whitespace_pattern": ...} | StructuredOutputsParams(whitespace_pattern=...) |
guided_decoding_backend | Remove it; the backend is a server setting | None |
response_format vs structured_outputs in vLLM: which one wins
At v0.30.0 both /v1/chat/completions and /v1/completions accept two structured-output surfaces.
The name inside json_schema is required. structured_outputs also carries modifiers (disable_any_whitespace, disable_additional_properties, whitespace_pattern). With the OpenAI Python SDK, anything vLLM-native goes in extra_body:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
schema = {"type": "object", "properties": {"total": {"type": "number"}}, "required": ["total"]}
# OpenAI shape
r = client.chat.completions.create(
model="extractor",
messages=[{"role": "user", "content": "Extract the invoice total."}],
response_format={"type": "json_schema", "json_schema": {"name": "invoice", "schema": schema}},
)
# vLLM-native shape
r = client.chat.completions.create(
model="extractor",
messages=[{"role": "user", "content": "Extract the invoice total."}],
extra_body={"structured_outputs": {"json": schema}},
)Four rules from the v0.30.0 source decide what actually happens when clients get creative.
response_format overrides. If a request sends both, structured_outputs_from_response_format() applies response_format on top of structured_outputs with a replace() helper. The same kind is overwritten; a different kind leaves two constraints set and is rejected by rule 3. A type: "text" response format leaves structured_outputs untouched.strict is decoration. The json_schema object accepts strict, but the conversion function never reads it. strict: true does not change enforcement. If you relied on it with a hosted API to reject extra keys, write "additionalProperties": false into the schema.json, regex and choice raises "You can only use one kind of constraints for structured outputs ('json', 'regex' or 'choice').", and other pairs fail the same check inside StructuredOutputsParams. On chat, a structured_outputs json, regex or choice constraint plus a named tool_choice is rejected as well; response_format plus a named tool_choice is not rejected, the tool's schema silently replaces it.guided_decoding_backend key is ignored like any other unknown field. The backend is a server setting.Rule 3 fails loudly, which is the good kind of failure. Rule 2, the same-kind override in rule 1 and the tool_choice replacement change behaviour without any error. Prompting still matters inside the constraint: the docs recommend describing the schema in the prompt, and our post on prompt structure for consistent JSON outputs covers that side.
| Surface | Shape | Constraint kinds |
|---|---|---|
response_format | {"type": "json_schema", "json_schema": {"name": ..., "schema": {...}}}, or json_object, structural_tag, text | JSON schema, any JSON object, structural tag |
structured_outputs | Top-level object with exactly one kind set | json, regex, choice, grammar, json_object, structural_tag |
Why structured output is not enforced with a vLLM reasoning parser
With a --reasoning-parser configured, vLLM must not constrain the model's thinking, so the grammar waits. How the parser splits output into reasoning and content is covered in our post on empty tool_calls and vLLM reasoning parsers. The structured-output gate is a separate code path, _get_constraint_start in vllm/v1/structured_output/__init__.py, and at v0.30.0 it works like this:
enable_in_reasoning is True, the grammar applies from the first generated token.reasoning_ended is initialised from is_reasoning_end(prompt_token_ids). While it is False, every step returns unconstrained until the parser sees the end of reasoning in the output. Once the grammar starts, the flag latches.On /v1/chat/completions the chat template and the model's think tags give the parser something to find. On /v1/completions there is no chat template. For qwen3, the parser scans the prompt for think tokens and, finding none, returns not wait_for_reasoning, and wait_for_reasoning follows enable_thinking, which defaults to True. The grammar never opens.
Issue #57726, reported on 0.27.1 with Qwen3-0.6B at temperature 0 and max_tokens 40, shows the result. Without a reasoning parser: {"name": "John Doe", "age": 30}, finish reason stop. With --reasoning-parser qwen3: an unclosed object with extra keys, finish reason length. choice and regex were unconstrained too. The chat endpoint on the same server behaved correctly with thinking on or off. A contributor confirmed the behaviour on a main commit on September 20, and the gate logic it describes is unchanged in the v0.30.0 source. PR #56200, which ships in v0.30.0, centralised how the end of reasoning is derived; it did not add an opener for completions requests.
The same gate produces a second failure when a model skips its think channel. #57714 reports that kimi_k3 never signals reasoning end without a think marker, so response_format, JSON schema and regex output all come back free-form. Fix PR #57718 was open and alternate PR #57758 closed unmerged as of September 30, 2026.
Three workarounds, in order of preference:
/v1/chat/completions. This is the path the gate is designed around.stop by appending <think>\n\n</think>\n\n to the prompt, which gives the parser the end marker it scans for.enable_in_reasoning. The docs document it for Qwen3 Coder models (v0.11.2+). It is server-wide, and it constrains from token 0, so every structured request on that server loses free-form reasoning:vllm serve Qwen/Qwen3-8B --reasoning-parser qwen3 \ --structured-outputs-config.enable_in_reasoning=True
If your workload depends on the model thinking before it answers, the third option trades that away. Use the first.
JSON schema features xgrammar does not enforce in vLLM
--structured-outputs-config.backend accepts auto, xgrammar, guidance, outlines and lm-format-enforcer, and defaults to auto. Its docstring warns that auto makes "opinionated choices" that are "subject to change in each release". At v0.30.0 auto tries xgrammar first and falls back only when xgrammar validation fails, either because has_xgrammar_unsupported_json_features flags the schema or because xgrammar cannot compile it: to guidance, or to outlines for non-tekken Mistral tokenizers or schemas with features guidance does not support.
The last two rows are silent. In #56556, on 0.28.1rc0 with xgrammar 0.2.3, a schema whose allOf combined a $ref base requiring x as an integer with minimum 10 and a branch requiring a string y accepted {"x":5,"y":"a"}, {"x":15} and even "hello". Guidance, running llguidance 1.7.6, enforced the same schema. The reporter points out that Pydantic and OpenAPI tooling emit multi-branch allOf for model inheritance, so this lands on ordinary extraction models. In #57550 a property declared as {"pattern": "^[a-z]+$", "maxLength": 3} with no type went undetected and xgrammar accepted {"code":"abcdefghij"}.
Fix the schema first, because that works on every backend. Flatten inherited models so each object has one properties block, and give every constrained property an explicit type:
{
"type": "object",
"properties": {
"code": {"type": "string", "pattern": "^[a-z]+$", "maxLength": 3}
},
"required": ["code"]
}With the explicit "type": "string", the pattern-plus-length check fires and auto routes the schema away from xgrammar. If you cannot change the schemas, for example because they are generated from shared Pydantic models, pinning guidance is the option the #56556 repro supports:
vllm serve Qwen/Qwen3-8B --structured-outputs-config.backend=guidance
Pinning any backend removes auto's rerouting, so run the canary below with your real production schemas after the change. Do not pin xgrammar as the "strict" choice. The branch that handles it is commented # xgrammar with no fallback: flagged schemas are rejected instead of rerouted, and the two silent gaps above stay silent.
vLLM is not alone. The llama.cpp grammars README lists its JSON-schema-to-grammar limitations and says "Unsupported features are skipped silently", including uniqueItems, contains, not and if/then/else. Ollama v0.34.4, released September 23, 2026, says only that structured outputs on thinking models now apply in a single pass and that its bundled XGrammar was updated.
| Schema feature | Flagged at v0.30.0? | Result under auto |
|---|---|---|
multipleOf on integer or number | Yes | Falls back |
uniqueItems, contains, minContains, maxContains | Yes | Falls back |
String format outside xgrammar's list | Yes | Falls back |
pattern or format with minLength/maxLength | Yes (xgrammar "silently drops" the lengths, per the source comment) | Falls back |
Complex propertyNames or patternProperties combinations | Yes | Falls back |
Multi-branch allOf | No (#56556) | xgrammar, constraint ignored |
pattern plus maxLength on a property with no type | No (#57550) | xgrammar, constraint ignored |
How to test that vLLM structured output is enforced
The check that catches a grammar that never engages is a request whose only legal output is a value the model would never write on its own. A single-property object with an enum of CANARY-7F3Q, sent with a prompt about the weather, has exactly one valid completion if the grammar is live. If it is not, the model writes about the weather. Send it on both endpoints through both surfaces, and assert three things: the text parses, it equals the expected object, and the finish reason is stop.
import json
import os
import sys
from openai import APIError, OpenAI
BASE_URL = os.environ.get("VLLM_URL", "http://localhost:8000/v1")
MODEL = os.environ.get("VLLM_MODEL", "extractor")
# Reasoning tokens count against max_tokens; leave room for thinking on reasoning models.
MAX_TOKENS = int(os.environ.get("CANARY_MAX_TOKENS", "1024"))
client = OpenAI(base_url=BASE_URL, api_key="-")
TOKEN = "CANARY-7F3Q"
SCHEMA = {
"type": "object",
"properties": {"status": {"type": "string", "enum": [TOKEN]}},
"required": ["status"],
"additionalProperties": False,
}
RESPONSE_FORMAT = {"type": "json_schema", "json_schema": {"name": "canary", "schema": SCHEMA}}
PROMPT = "Write one sentence about the weather." # unrelated on purpose
def chat(**kwargs):
r = client.chat.completions.create(
model=MODEL, messages=[{"role": "user", "content": PROMPT}], max_tokens=MAX_TOKENS, **kwargs
)
return r.choices[0].message.content, r.choices[0].finish_reason
def completion(**kwargs):
r = client.completions.create(model=MODEL, prompt=PROMPT, max_tokens=MAX_TOKENS, **kwargs)
return r.choices[0].text, r.choices[0].finish_reason
LEGS = {
"chat/response_format": lambda: chat(response_format=RESPONSE_FORMAT),
"chat/structured_outputs": lambda: chat(extra_body={"structured_outputs": {"json": SCHEMA}}),
"completions/response_format": lambda: completion(extra_body={"response_format": RESPONSE_FORMAT}),
"completions/structured_outputs": lambda: completion(extra_body={"structured_outputs": {"json": SCHEMA}}),
}
failed = 0
for name, call in LEGS.items():
try:
text, finish = call()
ok = json.loads(text or "") == {"status": TOKEN} and finish == "stop"
except (APIError, json.JSONDecodeError) as exc:
text, finish, ok = repr(exc), None, False
failed += not ok
print(f"{'PASS' if ok else 'FAIL'} {name} finish={finish} text={(text or '')[:80]!r}")
sys.exit(1 if failed else 0)Each leg maps to rows in the table. A qwen3-style reasoning parser on the server fails the completions legs, per #57726, which is the finding you want before a pipeline starts calling that endpoint. A reasoning model that skips its think channel, the #57714 case, can fail the chat legs. An HTTP 500 fails the leg instead of crashing the job, and a truncated or whitespace-padded answer fails on the parse or on finish_reason. The legacy completions endpoint of the OpenAI SDK has no response_format argument, which is why that leg passes it through extra_body.
The canary proves the grammar is live. It cannot prove that your production schemas avoid the allOf and type-less gaps, so add a second leg per production schema with a prompt that invites a violation, and validate the result against the schema. The canary itself cannot catch a client that still sends guided_json, because it builds its own requests. On v0.30.0 the server log can: after a deployment has taken real traffic, fail the check if the warning appears, remembering it is written one time per process for each distinct set of removed fields:
! grep -qF "removed guided-decoding field" /var/log/vllm/server.log
Run the canary in CI against a pinned image, and again after every upgrade and every change to --reasoning-parser, --structured-outputs-config or speculative-decoding settings.
On-prem extraction: a silent 200 is a data-integrity incident
Regulated extraction pipelines (claims intake, KYC documents, lab reports, invoices) write model output into systems of record. When that pipeline trusts HTTP 200 as proof the schema held, every silent row in the table becomes unvalidated free text flowing into a record someone will audit. That is a data-integrity incident with a paper trail, not a parse error.
Running vLLM inside your own perimeter changes the answer in three ways.
Upgrades arrive in large jumps. An air-gapped or change-controlled environment that upgrades on a slow cycle can cross several minor versions in one step. A jump from 0.11.x to a current release is exactly how a guided_json caller loses its constraint with no error. Our guide to air-gapped vLLM deployment covers staging the image. Put the guided_ grep and the canary into the change ticket for every vLLM version bump.
The warning lands where nobody looks. The v0.30.0 warning goes to the server log once per process for each distinct set of removed fields. If those logs ship to a SIEM the ML team does not watch, nobody sees it. Add an explicit alert on the string "removed guided-decoding field".
You own the knobs a hosted API would own. Engine version, backend, reasoning parser and speculative decoding are your configuration, and each one appears in the table above. Treat the server as production infrastructure with a pinned version, the posture described in our post on hardening a vLLM inference server. Validate against the schema at the pipeline boundary, quarantine failures for review, never coerce them into shape, and make the canary a deployment gate. More on running open models in this setting is in the LLM models pillar.
This week, run the guided_ grep across every repository that calls your vLLM endpoint, then run the four-leg canary against the server you have in production today. If any leg fails, you already know which row of the table to read.
FAQ
Quick answers to the questions this post tends to raise.



