Same bytes, different parser

/ Article
[ Fig. 1 ]

This week I sat down with two copies of the same blog posts: the MDX files that actually publish, and HTML twins a reviewer reads so they can see the live page without clicking through the site. MDX is Markdown with components mixed in. A twin is the same article rendered as HTML so a person can review it like a webpage.

The Tech captions were gone.

If you don’t care about parsers, the takeaway is this: identical bytes are not an identical document. The format you paste into is not the format you tested.

What a Tech label is

Some of my posts put a one-word caption over a command. On the live page it looks like a tiny heading that says Tech, then the command itself. In the source it was often just four spaces at the start of a line, the CommonMark way of saying “this is a code block.”

CommonMark is the usual Markdown contract. Four spaces of indent means code. I had been using that indent as a cheap caption: indent the word Tech, and the renderer draws a little code-looking label.

It worked in the HTML twins. It did not work in the files that ship.

The wrong turn

The first theory was that the twin generator was stripping tags. There is a pass that pulls HTML components out of the body so the twin stays plain markdown. A caption wrapped in a <Tech> tag would vanish if that pass ate the tag and forgot to keep the word.

That theory lasted until I looked at files with no tags at all. Just an indented line. Those captions were gone too.

The second theory was CSS. Maybe the label was in the HTML and the stylesheet was hiding it. I searched the twin output. The word was not there. You cannot hide a node that was never written.

Same file, two parsers

Here is the part that is easy to miss when you copy a post from one tree to another.

In CommonMark, this is a code block:

    Tech

Four spaces, then the word. The HTML twin’s extractor had a shield for that: if a line looks like indented code, do not treat it as a tag to strip, because it is a sample, not markup. That shield is how you keep a tutorial from eating its own examples.

MDX does not treat a four-space indent as code. In MDX, a fenced block (the one with backticks) is code. An indent is just an indent.

So the shield was looking at leftover spaces in front of a <Tech> tag, deciding “this is a code sample, leave it alone,” and never handing the tag to the caption extractor. The caption dropped. The command still published. The label did not.

I had been reviewing the HTML and blaming the twin. The twin was telling the truth about what MDX actually parsed.

The bytes were the same. The grammar was not.

What I changed

I stopped asking the indent to mean “caption.”

The labels stay. They are headings now, or a short summary line above the command, which both parsers agree on. I did not add a special MDX plugin to teach MDX CommonMark’s indent rules. Teaching a second parser to imitate the first is how you get a third bug next month.

The shield stays too, for real code samples in the HTML twins. It just does not run on MDX bodies, because those bodies were never using indent as code. An unused CSS token I almost deleted in the same pass is a different story: unused is not the same as wrong. The indent was wrong.

The named thing

Same bytes, different parser.

I keep wanting a file compare to be the whole test. If the strings match, the documents match. That is true for a compiler that has one grammar. It is false the moment you have Markdown and MDX in the same workflow, or CommonMark and GitHub Flavored Markdown, or “it looks fine in the editor preview” and “it looks fine on the host.”

The preview you trusted was a different program than the publisher.

This is the same class of bug as a test suite that encodes the author’s mental model and then calls the result coverage. The shield was correct for CommonMark. It was a lie for MDX. Green output in the twin pipeline did not mean the caption survived. It meant the shield had classified the line and moved on.

What I watch for now

When I copy a post into a new format, I do not ask “did the text survive?” I ask “did the meaning that depended on whitespace survive?” Captions, nested lists, indented samples, anything that is a signal in one grammar and decoration in another.

If a label matters, I make it a heading. Headings are boring. They also survive a paste.

Related: The pipeline is the review, My passing tests encoded the fail-open bug as correct behavior, The outdated comment that wouldn’t die, and AI code needs a receipt.