A chat product often needs more than the visible reply: citations, related queries, a follow-up form. Structured output looks like the cheap way to get them: one schema, one completion, then read reply and ship it. But the answer now lives inside a string field. It is like writing an email inside a spreadsheet cell: same words, but every line break and quote must be escaped. Sometimes a second document is nested inside the first.
Prior work disagrees on who pays. Tam et al. (opens in a new tab) blamed format restrictions, .txt (opens in a new tab) blamed prompting, and The Format Tax (opens in a new tab) found 92% of the loss already came from asking for a format. Our own first run, with five models and one decoder-enforced JSON arm, blamed the decoder. That read was too simple.
So we separated the effects. We tested freeform, JSON asked in the prompt, JSON enforced by the decoder, YAML, XML, and a two-turn reformat, on nine models and five benchmarks. The answer is not a single tax: it depends what the string is asked to hold. Reasoning, and code inside JSON, mostly survive. Instruction following drops on every single-call format, and replies lose most of their markdown. A second call that wraps a freeform draft recovers most of the instruction-following drop (8 of 9 models within 5 points).
Six ways to wrap the same reply
aider (opens in a new tab) found the same tax on code returned inside JSON. Parikh (opens in a new tab) measured a register gradient at the token level on single-word answers: asking for JSON costs 0.22 bits of surprisal, XML 0.19, YAML and CSV close to zero. This run separates format asked in the prompt from format enforced by the decoder, on the same model and the same message.
Three questions, three arms each answers:
| Question | Arm(s) |
|---|---|
| Does asking for a format cost anything, with no decoder involved? | json_prompted, yaml_prompted, xml_prompted |
| Does the decoder constraint cost more on top of that? | json_constrained vs json_prompted |
| Does moving the format to a second call recover the loss? | two_turn |
The envelope is this schema, in every format:
class AIReply(BaseModel):
reply: str = Field(..., min_length=1, description=REPLY_DESCRIPTION)
related_queries: list[str] = Field(
default_factory=list,
max_length=3,
description="Up to 3 short follow-up queries the user might "
"plausibly ask next. Never put your answer here.",
)REPLY_DESCRIPTION itself is a variable we test on one benchmark (see below), not a fixed string: run 1 always described reply as a chat response, which turns out to matter a lot once the task wants a JSON document back.
The same answer, five ways
One GSM8K item, one model (gpt-5.6-luna), same question, five arms. This is what the harness actually sends to the scorer versus what it throws away:
Freeform, scored as-is:
Eliza earns regular pay for the first 40 hours:
40 × $10 = $400
Her overtime rate is 1.2 × $10 = $12 per hour
...
Total earnings: $460JSON, decoder-enforced (json_constrained), the whole object is the API response, reply is deserialized before scoring:
{"reply":"Eliza earns regular pay for the first 40 hours:\n40 × $10 = $400\n...\nTotal earnings: $460","related_queries":[]}JSON, prompt-only (json_prompted), same shape, no decoder constraint, the model can fail to close a brace and often does not:
{"reply":"Eliza earns regular pay for the first 40 hours:\n40 × $10 = $400\n...","related_queries":[]}YAML, prompt-only (yaml_prompted):
reply: |
Eliza earns $10 per hour for the first 40 hours:
40 × $10 = $400
...
#### 460
related_queries:
- How is overtime pay calculated?XML, prompt-only (xml_prompted):
<ai_reply>
<reply>Eliza earns $10 per hour for the first 40 hours:
40 × $10 = $400.
...
#### 460</reply>
<related_queries>
<query>How much would Eliza earn if she worked 50 hours?</query>
</related_queries>
</ai_reply>Five formats, one underlying answer. IFEval and JSONSchemaBench Hard below are about what happens when that answer has to carry stricter content: exact wording, or a second document.
Nine models, five benchmarks
Nine models, one seed, n=100 per task (80 on MT-Bench, its full loaded split), one sample per item. with_structured_output(..., method="json_schema") for json_constrained; LangChain's own choice of native path per provider (forced tool call on Anthropic, native schema elsewhere).
| Model | Provider |
|---|---|
| GPT-5.6 Sol, GPT-5.6 Luna | OpenAI |
| Claude Sonnet 5, Claude Haiku 4.5 | Anthropic |
| Gemini 3.8 Flash, Gemini 2.5 Pro | |
| Grok 4.6 | xAI |
| Mistral Medium 3.5 | Mistral |
| Kimi K3 | Fireworks (Moonshot, open-weight) |
One benchmark could have shown a false floor or a false ceiling. GSM8K alone would say the envelope is free. IFEval alone would say it is costly everywhere. Five tasks, each asking a different question about the same wrapping:
| Task | What it checks | Why this one |
|---|---|---|
| GSM8K | A checkable number at the end of a written solution | Does wrapping make the model more or less able to reason, when the answer is one checkable fact |
| HumanEval | Executable Python inside the reply string | Does code inside the reply string still run? |
| IFEval | Exact surface-form instructions ("no commas", "use bullets") | How much concentrating on the envelope hurts following other instructions |
| JSONSchemaBench Hard | A second JSON document, nested inside reply | Nested JSON inside a JSON string: the model must escape newlines, quotes, and the rest, or the inner document is invalid |
| MT-Bench | An open chat reply, no task-level format | Free writing. Compare constrained vs unconstrained replies on quality (an LLM judge) and on form (length, markdown, emoji) |
MT-Bench is judged pairwise (each arm versus freeform, position randomized) by gpt-5.6-terra on all pairs and claude-opus-5 on a 20% subsample, neither in the 9-model roster, so no model judges its own family. The judge sees markup-stripped text. Length, markdown markers, and emoji are scored separately on the raw reply. Every other task uses an exact-match or execution-based scorer, not an LLM judge.
Instruction following drops on every prompted or constrained format
Every model loses IFEval accuracy on every prompted or constrained single-call format, including decoder-enforced JSON. two_turn is not the worst IFEval arm on any model. A second call that is told to copy the draft, and still sees the original user request, sits within 5 points of freeform for 8 of 9 models. Mistral is the exception (-8 points; the 95% CI excludes zero).
Turn 2 can still redo the task, so this is not a pure wrap. IFEval scores strict surface form ("write exactly N sentences", "no commas"), so a remaining drop is still wrapping fighting the instruction.
If your product depends on exact instructions, write freeform first and wrap it in a second call, at the cost of latency.
JSONSchema Hard: the decoder was not the main problem
JSONSchemaBench Hard is the nested case: the API already forces AIReply, and the task wants a second JSON document inside reply. Run 1 tested only decoder-enforced JSON here and saw a near-total collapse on 4 of 5 models. This run splits prompted JSON from constrained JSON, and adds YAML and XML:
Decoder-enforced JSON is close to flat for 7 of 9 models. The two exceptions are Gemini 2.5 Pro (-79) and Haiku (-11). Prompted formats are where most of the rest of the grid breaks:
json_prompted: Haiku -41, Mistral -33, Gemini 2.5 Pro -32, Sonnet -16.yaml_prompted: Mistral -67, Sonnet -28.xml_prompted: Haiku -52.
YAML and XML costing this much is the opposite of what Parikh found for single-word answers, where YAML cost nothing.
two_turn is near freeform on most models (Mistral -3, Kimi +4). Gemini 2.5 Pro is +8 points (the 95% CI excludes zero): turn 2 still sees the original request, so it can write the nested document instead of wrapping a chat-framed draft.
Constrained decoding does what it says: it guarantees the outer envelope. It does nothing about the document inside the string, and on most models that turns out to be fine. Without the decoder holding the outer shape, a document inside a document is where models actually struggle.
Gemini 2.5 Pro is the one model where json_constrained still collapses hard (-79 points). That number turned out to depend entirely on one sentence.
One sentence in the schema: 0% versus 87%
AIReply.reply's field description is either chat-framed ("your main natural language response to the user") or task-framed ("the complete output for the request; if a JSON value is asked for, put exactly that here"). Same schema, same decoder, same model, same item, JSONSchemaBench Hard, json_constrained:
// chat-framed reply description, item jsonschema-o1052
{"reply": "I've created a sample JSON object based on the `people.aarhusteater` schema. It includes details for a fictional person...", "related_queries": []}// task-framed reply description, same item, same model
{"reply": "{\n \"uuid\": \"a1b2c3d4-e5f6-1234-a456-426614174000\",\n \"id\": \"123456789\",\n \"domain\": \"people\",\n ...", "related_queries": []}| Model | Chat-framed compliant | Task-framed compliant |
|---|---|---|
| Gemini 2.5 Pro | 0% | 87% |
| Claude Sonnet 5 | 88% | 89% |
| Claude Haiku 4.5 | 77% | 76% |
| GPT-5.6 Sol | 97% | 100% |
| Grok 4.6 | 98% | 98% |
Every chat-framed reply from Gemini 2.5 Pro on this cell starts as valid JSON already. The model is not confused about syntax. It narrates ("I've created a sample JSON object...") instead of returning the document, because the field description told it to. Every other model absorbs the same description with a much smaller penalty. The field description is not a footnote here. On at least one model, it is the whole effect.
If your schema has a reply field, describe it as the task output, not as a chat response.
HumanEval: constrained JSON mostly survives, YAML and XML often do not
json_constrained is flat on 8 of 9 models. The one exception, Mistral (-93), is not an escaping failure: on every one of those 100 items the model writes a paragraph explaining what it would do instead of emitting code, a chat-framed refusal to just be code, the same failure mode as the JSONSchema ablation above. two_turn stays within 4 points of freeform. Luna is the one cell whose 95% CI excludes zero (-4).
yaml_prompted and xml_prompted are the real HumanEval story:
yaml_promptedbreaks code for Kimi K3 (-97), Grok 4.6 (-87), Haiku (-83), Sol (-78), and Sonnet (-62).xml_promptedbreaks code for Grok 4.6 (-93), and Haiku'sjson_promptedcell drops too (-83).
This reframes what "code stays flat under structured output" actually means: it held for the JSON envelope, not for YAML or XML carrying code. Whether the mechanism is mostly the model narrating instead of coding, or code broken by YAML/XML escaping, is not separated in this run: both patterns show up in a manual read of the failures, in different proportions per model.
If the reply carries code, do not ask for YAML or XML.
GSM8K stays flat
A checkable final number, wrapped in almost any format, survives. This matches the .txt replication more than Tam et al.'s original finding, and confirms the earlier run's read: reasoning to a number is the easiest case here: 6 of 9 models stay within 7 points on their worst arm.
The MT-Bench judge does not see a broad quality collapse
Most cells sit close to 50%: the judge is not detecting a broad quality collapse the way the IFEval and JSONSchema numbers might suggest. Mistral is the sharpest drop, losing on json_constrained (30%) and yaml_prompted (30%), the same two arms where its other metrics also broke down. YAML is the next place the judge notices: Sol 33%, Luna 33%, Kimi 36%, Gemini 2.5 Pro 40%. Haiku XML goes the other way (67%). For the remaining cells, an envelope changes measurable structure (instruction compliance, schema compliance, code correctness) well before it changes what a reader-side judge calls a worse reply.
The judge is also blind to form on purpose. Both sides of every pair are markup-stripped before scoring, because LLM judges overweight markdown. The next three charts measure the thing the judge cannot see: how the same reply looks.
The envelope flattens markdown
Every prompted or constrained single-call arm loses markdown on every model. The drop is large: Luna XML -91%, Grok YAML -97%, Mistral constrained JSON -100% (zero markers in 80 replies). Constrained JSON is mild for some models (Sol -13%, Haiku -13%) and not for others (Gemini 3.8 Flash -80%, Gemini 2.5 Pro -63%, Mistral -100%). two_turn stays near the freeform markdown baseline on several models (Grok +23%, Sol +17%). Mistral is still -44% after the second call, so turn 2 is not a mechanical copy.
Grok barely uses markdown even in freeform (2.0 markers per reply), so its percentage drops sit on a tiny baseline. Sol starts at 12 markers. The other seven models start from 13-20.
Replies get shorter, with two exceptions
Prompted YAML and JSON usually shorten the answer, by about 10-55%. The largest cuts are Mistral YAML (-56%) and Gemini 2.5 Pro constrained JSON (-55%) and YAML (-53%). two_turn stays close to freeform on every model. That is expected if turn 2 mostly copies, but this run does not prove a faithful wrap.
Constrained JSON is mixed, not uniformly shorter. Haiku's replies get longer (+23%), Sonnet's slightly longer (+9%), Sol almost flat (-5%). The "envelope always truncates" story does not survive that split. What does survive: the prompted formats, YAML especially, cost length as well as markdown.
Emoji was already rare, and gets rarer
Emoji is a minority habit. Haiku uses it in 10% of freeform replies, Sonnet in 9%, Kimi in 6%. Those rates fall to about 0-1% on YAML and prompted JSON, and recover toward the baseline on two_turn. Kimi's constrained JSON arm is the exception that stays at 6%. This is a small count (8, 7, and 5 emoji-using freeform replies out of 80), so treat it as a direction, not a precise tax. The markdown collapse above is the form result that is large enough to take at face value.
The safety net held
None of this would mean much if json_constrained had silently degraded to plain JSON mode or text on some model. It did not: of 5 220 json_constrained calls across all 9 models, 5 219 resolved to their expected native path (forced tool call on Anthropic, native schema elsewhere). One cell, on Mistral, fell back to unconstrained text.
What this does not prove
This is one envelope (AIReply.reply plus related_queries), one LangChain integration per provider, and two reply descriptions tested together only on one benchmark. It does not test tool-calling schemas, grammar-constrained custom tools (OpenAI's apply_patch pattern, Cursor's own patch grammar), or a task-framed description tested across every benchmark. Kimi K3 is the only open-weight model in the roster, served through Fireworks' own stack rather than reference weights, so this run cannot say whether open-weight models pay a larger format tax as a group, only report where this one model landed.
Both turns of two_turn used the same model. The product case is a size split: a capable model writes freeform, a smaller model only wraps. This run does not test that pair. Two sequential calls also add latency before the envelope exists. Turn 2 still sees the original user request, so it can redo the task rather than only copy the draft.
The reading in the conclusion (code as a familiar tool-content shape, JSON training as extracted values) is a hypothesis about why the split looks this way. This run did not inspect training data or tool traces.
n=100 (n=80 on MT-Bench) is a fixed-seed random sample, not the full dataset. HumanEval and JSONSchemaBench Hard failures were not separated by cause (prose refusal versus broken serialization) beyond the one ablation cell above; a model-by-model read of every failure would be needed to split the two.
MT-Bench's judge decision is pairwise preference, not an absolute quality score, and covers one turn. The judge scores markup-stripped text, so Figures 7-9 (markdown markers, word count, emoji) are a separate measurement on the same reply strings, not something the judge was asked. Emoji rates sit on a small count (at most 8 of 80 freeform replies), so that chart is a direction, not a precise tax.
So, does enforcing JSON make answers worse?
Does enforcing JSON make answers worse? Not in one direction. It depends what the reply string is asked to hold.
A checkable reasoning answer mostly survives. GSM8K stays flat once the number lives inside a JSON field. A fair reading: the decoder is not squeezing the chain of thought, only where the final number sits.
Nested JSON is the opposite case. The task is a document inside a document: double serialization, quotes and newlines escaped. That is where prompted formats break. Decoder-enforced JSON mostly holds the outer envelope (7 of 9 models within 5 points of freeform on JSONSchemaBench Hard). The inner document is where models struggle, and on Gemini 2.5 Pro one sentence in the field description was the whole effect (0% versus 87%).
Code is closer to GSM8K than to nested JSON. HumanEval mostly survives constrained JSON. A fair reading, not a measurement of training data: models already write code into tool content fields, so a string attribute that should contain a program is a familiar shape. YAML and XML are not, and they often wipe the code out. Mistral is the exception even under JSON: it writes a paragraph about the code instead of the code.
Open writing is where the envelope still shows. An LLM judge, reading markup-stripped text, does not call the enveloped replies worse as a class. The same strings are shorter and lose headings, bold, and lists. That also matches what JSON usually looks like in supervised data: extracted values, flattened and short, not formatted prose. The model is doing what that format looks like. We did not inspect a training corpus, so this is a reading of the form drop, not a proof of its cause.
Instruction following is the cost that survives all of the above. IFEval drops on every prompted or constrained single-call format. A same-model second call that first writes freeform recovers most of that drop (8 of 9 models within 5 points of freeform; Mistral -8). That recovery is not a mixed-size pipeline. Both turns used the same model, turn 2 still sees the original request so it can rewrite, and two sequential calls add latency.
We already moved the product off {reply, questions} toward a freeform message plus a dedicated ask-questions path. The envelope is simpler for the backend. It does not look like it makes the model less able to reason. It does flatten writing and fight extra instructions, and fixing the field description does not make those costs disappear.
Going forward
- Test the size split: a capable model writes freeform, a smaller model only wraps.
- Test the task-framed
replydescription on every benchmark, not only JSONSchemaBench Hard. - Test tool-calling schemas and grammar-constrained custom tools.