Is HTML "Strictly Better" Than Markdown for Claude Code?
Thariq Shihipar of Claude Code team fame posted an article last week, "The Unreasonable Effectiveness of HTML", arguing that Claude should output HTML files by default for basically everything. PR reviews, postmortems, status reports, the lot. It's a good piece, racked up 8.3M views, and most of it I agree with. But in the replies, when somebody pushed back to say markdown's fine for text-heavy stuff, Thariq held the line: "idk I think HTML is strictly better for all of that too." So I ran it all. Three of his use cases, instrumented end to end on Opus 4.7. What follows is what I measured, what surprised me, and where I ended up agreeing and disagreeing. Short version: he's right about half of it, and that half is very right. The other half I don't think holds up, and the data here is why.
TL;DR
Thariq's piece is good. Read it before you read this.
- Visual tasks (design exploration, prototypes, drag-and-drop widgets): HTML wins outright. Markdown can describe a design; HTML can show one. The 2-4x token cost is the price of admission, not a tradeoff.
- Text-with-structure tasks (PR reviews, technical explainers): both formats produce a real artifact. HTML adds polish for about 1.4-1.5x the dollar cost. In one of my two text cases (the rate-limiter explainer) the markdown surfaced more gotchas (12 vs 8) at lower cost. Which cuts against "strictly better."
- Does Claude read HTML better than markdown? (the question Thariq's argument quietly assumes a HTML-favorable answer to): two blinded judges scored Claude 7/7 on factual accuracy from both formats, identically. On explanatory specificity HTML got a ~7% edge, at 33% higher cost per call. Pick your poison.
- The "1M context handles it" hand-wave: at N=50 artifacts re-ingested K=3 times with cache evictions, HTML's overhead eats roughly the whole 1M window. Not nothing, but depending on the work being done maybe worth it.
Sample is small (N=1 per condition for content quality, three instrumented cases) and the methodology I used actually biases against markdown. Even with the thumb on the scale, "strictly better" isn't what I saw.
What Thariq actually said
He walks through five buckets of tasks where he reckons HTML should be the default. Roughly:
- Specs, planning, exploration (e.g. design exploration)
- Code review & understanding (e.g. PR review with annotated diff)
- Design & prototypes (e.g. an interactive checkout button)
- Reports, research, learning (e.g. a rate limiter explainer)
- Custom editing interfaces (e.g. drag-and-drop Linear ticket reorderer)
He flags three costs honestly: more tokens, slower generation, noisier diffs. He concludes the advantages outweigh them. Reasonable people can disagree on the conclusion.
What I want to look at is the reply to the comment. "Strictly better" is a specific, testable claim. That's the one I want to examine.
There are actually two questions here
The article treats format choice as a knob: pay a bit extra in tokens, get nicer output, decide if it's worth it. That framing works for half the use cases and quietly stops working for the other half. Here's the split.
Sometimes HTML can do something markdown structurally can't. Render a six-mockup grid of UI designs. Run a slider that updates a preview in real time. Drag a card across a kanban. Markdown isn't a worse version of these; there is no markdown version. Asking "is the extra cost worth it?" is like asking whether the airfare to Mars is good value when the alternative is staying home and looking at a globe.
Other times both formats can do the job, and HTML just costs more. PR reviews. Postmortems. Technical explainers. Implementation plans. Markdown handles all of them. HTML renders them more attractively. The substance lives in either.
The "use HTML for everything" framing treats both questions like the second one. The case for HTML gets stronger, not weaker, when you split them apart.
How I ran this
I picked three of Thariq's five use cases and instrumented them properly. Same prompt, both formats, capture every token. The other two are the categorical ones from the previous section; there's nothing fair to compare them against, so I just demo them at the end.
The three I instrumented:
1. Design exploration: onboarding screen, six approaches in a grid (his specs/planning bucket)
2. PR review of a real commit: toks commit c6d70f9, "Respect .gitignore when target has no .git directory"
3. Rate limiter explainer over the slowapi source: about 12K tokens of Python as substrate
The two I just demo: 4. Checkout button prototype (interactive sliders, animation, copy-as-prompt) 5. Linear ticket reorderer (drag-and-drop kanban with copy-as-markdown export)
For each instrumented case I ran three probes: generate the artifact, feed it back into a fresh agent session as context for a downstream task, and apply a semantic edit while capturing the diff. The first probe is the obvious one. The other two matter because that's where the cumulative token costs actually bite. Re-ingestion and editing happen constantly in real Claude Code workflows.
On the design exploration case I ran two methodologies side-by-side: Thariq's prompt unchanged versus a neutrally-phrased version of the same task. Why? Because his prompt literally says "create an HTML file." If you only do a "HTML" → "markdown" word-swap, you're asking Claude to make a markdown file in the shape of an HTML file, which is not the same as asking for markdown. Tracking? The difference between the two methodologies turned out to be larger than I expected: a 3.6x cost ratio under ablation, 2.1x under neutral phrasing. Worth being honest about.
What "artifact tokens" means below. I ran
toks <file> --for claudeon each generated artifact, using Anthropic's own tokenizer to count the file the way Claude would see it if fed back as context. That number is different fromusage.output_tokens(what the API reported as the generation length) and fromcache_creation_input_tokens(artifact + system prompt + user message, all of it). The first matters for re-ingestion cost. The other two matter for other things.
Where this is honestly a bit thin: - My reader-side reactions are N=1. One user (me), one session, I knew which file was which. No blinding, no randomization. Take them as gut reactions, not data. - Generation runs are also N=1 per condition. Content-quality findings like "markdown surfaced more gotchas" might be variance, not signal. K=5 per condition would tell us. I didn't run K=5. - Three instrumented cases is a sample, not a census. - The agent-as-reader test below covers shared-content questions only, i.e. facts present in both artifacts. I didn't test whether HTML's verbose-priming or callout-salience changes output quality (beyond binary accuracy) in ways that matter downstream. - I didn't test "colleagues actually read HTML more." That'd need users I don't have. - Methodology asymmetry is load-bearing, and I'm flagging it here so the reader knows the thumb is on the scale before the data tables start. I ran two methodologies only on the design exploration case. The PR review and rate-limiter cases used Thariq's verbatim prompt structure, which is HTML-affording ("create an HTML artifact..."). That biases against the markdown artifact compared to a neutrally-phrased prompt. The cost ratios I report for those cases (1.51x for PR review, 1.44x for rate limiter) would likely shrink under a neutral prompt; how much, I didn't measure.
Design exploration: this is where HTML earns its keep
Thariq's prompt (verbatim): "I'm not sure what direction to take the onboarding screen. Generate 6 distinctly different approaches, varying layout, tone, and density, and lay them out as a single HTML file in a grid so I can compare them side by side. Label each with the tradeoff it's making."
Cost data (verbatim methodology):
| Format | Artifact tokens | Output tokens | Generation time | Cost |
|---|---|---|---|---|
| HTML | 6,877 | 10,135 | 117 s | $0.51 |
| MD | 1,896 | 3,284 | 49 s | $0.27 |
Cost data (format-neutral methodology):
| Format | Artifact tokens | Output tokens | Generation time | Cost |
|---|---|---|---|---|
| HTML | 6,871 | 9,161 | 102 s | $0.48 |
| MD | 3,332 | 4,769 | 73 s | $0.34 |
Reading both artifacts side by side, my reaction was unambiguous:
"HTML is massively nicer to read in these cases versus the markdown. It's much better ... it's not speed that makes the entire difference. It's quality and detail of visual output. This question is fundamentally visual and HTML allows for near-perfect visual communication of the ideas vs text-only in markdown."
The HTML rendered six high-fidelity onboarding mockups in iframes inside a responsive grid. The markdown produced a six-section text document that described what each mockup would look like.
The HTML artifact, rendered above. Six approaches, side by side. This is the thing markdown structurally can't do.
Three things worth noting from this case:
- Markdown leaks HTML, but only when the prompt corners it. When I ran Thariq's verbatim prompt (with its "single markdown file in a grid" phrasing), the model embedded 10 HTML tags into the "markdown" output (
<table>,<tr>,<td>). Of course it did; there's no markdown way to lay out a grid. When I dropped "in a grid" from the prompt, the markdown came out clean. Zero HTML tags. So if you've ever wondered why your markdown agents sometimes spit out HTML, here's one answer: you asked them to do something markdown can't do. - Methodology shifts the cost ratio. Verbatim ablation: 3.63x artifact tokens; format-neutral: 2.06x. Single-number ratios reported without methodology disclosure should be read skeptically.
- Generation time partially supports Thariq's "2-4x slower" admission. Measured 1.4-2.4x depending on methodology. The high end of his range is reachable, the low end is below what I measured.
The call: HTML wins outright. The token cost isn't a tradeoff, it's just what the capability costs.
PR review: a fair fight, and HTML costs about 40% more for the polish
Methodology note: this case used verbatim ablation only (the format-neutral methodology was only run on the design exploration). The verbatim prompt structure ("create an HTML artifact...") is HTML-affording and biases against markdown. The 1.51x cost ratio reported here would likely shrink under a neutrally-phrased prompt; I didn't measure by how much.
Thariq's prompt (lightly adapted for my substrate): "Help me review this PR by creating an HTML artifact that describes it. I'm not familiar with the gitignore/path-resolution logic so brace on that. Render the actual diff with inline margin annotations, color-code findings by severity..."
Substrate: toks commit c6d70f9, a real bug fix (~30 lines of code change plus a new test).
Cost data:
| Format | Artifact tokens | Output tokens | Generation time | Cost |
|---|---|---|---|---|
| HTML | 11,223 | 19,308 | 234 s | $0.82 |
| MD | 3,810 | 10,437 | 154 s | $0.54 |
Ratios: 2.95x artifact tokens, 1.85x output tokens, 1.52x time, 1.51x cost.
The markdown PR review is, frankly, a real PR review. TL;DR verdict at the top ("Approve with two non-blocking notes"), a severity legend in a markdown table, a structured walkthrough of the gitignore logic, line-by-line diff annotations, a test-coverage assessment. If a teammate sent me this in a Slack DM I'd be perfectly happy.
HTML adds a color-coded verdict bar, rendered diffs with green/red highlighting and line numbers, eight severity-tagged finding cards (Pass / Concern / Nit), and inline severity callouts on the diff. Prettier. More navigable. The substance is the same.
The HTML PR review. Verdict bar, severity-tagged findings, annotated diff. Polished. But the markdown version below has the same substance.
This review found: 0 CRITICAL · 0 HIGH · 2 MEDIUM · 3 LOW · 2 POSITIVE.
Bracing the gitignore / path-resolution logic
I want to walk this carefully because the bug is exactly the kind of thing that hides in path semantics. Here's the model after the change:
target = the directory the user pointed toks at
git_root = nearest ancestor containing .git (or None)
gitignore_root = git_root if git_root else target # NEW
gitignore_root is then used for two things:
- As the walk root for
.gitignorediscovery —load_gitignore_specs(git_root=...)doesos.walk(git_root)and concatenates patterns from every.gitignoreit finds, prefixed by their relative path under that root. - As the basis for relativization at match time —
file_path.relative_to(gitignore_root)produces the path string passed topathspec.match_file.
These two uses must use the same root for matching to be correct. The PR keeps them in lockstep — that's the load-bearing invariant, and it's preserved.
Walking the four cases
| Case | git found? | git_root |
gitignore_root (after) |
Discovery walks | Relativizes against | Correct? |
|---|---|---|---|---|---|---|
| Target inside a git repo | yes | /repo |
/repo |
/repo |
/repo |
yes (unchanged) |
| Target IS the git root | yes | target |
target |
target |
target |
yes (unchanged) |
Target has its own .gitignore, no .git anywhere |
no | None |
target |
target |
target |
yes — newly fixed |
Target has no .gitignore, no .git |
no | None |
target |
target (yields no patterns) |
target |
yes — load_gitignore_specs returns None, so the guard if gitignore_spec and gitignore_root: short-circuits |
The fourth row is the one I'd want a reviewer to verify by reading. load_gitignore_specs returns None when patterns is empty, the assignment becomes gitignore_spec = None, and the per-file guard skips matching. No regression.
What the fallback does NOT do
There's one behavior it would be easy to assume but isn't true: the fallback does not search upward for .gitignore files when there's no .git. If the layout is
/parent/
.gitignore # contains "*.log"
child/
debug.log
…and the user runs toks /parent/child, debug.log will be scanned. find_git_root returns None (no ancestor has .git), gitignore_root falls back to child, and os.walk(child) never sees /parent/.gitignore.
This is consistent with the commit message ("fall back to the target directory itself") and consistent with how git itself behaves (no .git → no project boundary → no upward .gitignore chain). I'd just like one line in the docstring saying so, because the next person to think about this will think about it again otherwise.
The diff, annotated
Below: each hunk, followed by per-line notes. Annotations cite line numbers from the new file.
Hunk 1 — scanner.py lines 104-112
target = target.resolve()
if not target.is_dir():
raise ValueError(f"Not a directory: {target}")
gitignore_spec = None
- git_root = None
+ gitignore_root = None # ← (A)
if not no_gitignore:
git_root = find_git_root(start=target)
- if git_root:
- gitignore_spec = load_gitignore_specs(git_root=git_root, target=target)
+ gitignore_root = git_root if git_root else target # ← (B)
+ gitignore_spec = load_gitignore_specs(git_root=gitignore_root, target=target) # ← (C)
| Mark | Annotation | Severity |
|---|---|---|
| A | Rename is the right call — this variable now means "the root we relativize against," not "the git repo root." Name follows semantics. | [POSITIVE] |
| B | Ternary is readable. Could equivalently be gitignore_root = git_root or target, which is shorter and idiomatic Python. Style preference; either is fine. |
[LOW] |
| C | The target=target keyword arg is now unused inside load_gitignore_specs — target_resolved = target.resolve() is computed and never read. Pre-existing dead code, but this PR is the natural moment to either delete the parameter or use it. See finding F-3 below. |
[MEDIUM] |
Hunk 2 — scanner.py lines 128-135
if file_path.is_symlink() and file_path.is_dir():
continue
- if gitignore_spec and git_root:
- rel = file_path.relative_to(git_root)
+ if gitignore_spec and gitignore_root: # ← (D)
+ rel = file_path.relative_to(gitignore_root) # ← (E)
if gitignore_spec.match_file(str(rel)):
continue
| Mark | Annotation | Severity |
|---|---|---|
| D | Guard updated in lockstep with the rename. The two checks (gitignore_spec, gitignore_root) are now both truthy iff we successfully built a spec — gitignore_spec alone is sufficient (since spec is only built when root is set), but the redundant guard is defensive and harmless. |
[POSITIVE] |
| E | relative_to(gitignore_root) is safe: file_path comes from os.walk(target), and gitignore_root is either target or an ancestor of it, so file_path is always under gitignore_root. No ValueError risk. |
[POSITIVE] |
Hunk 3 — tests/test_scanner.py lines 99-110
+ def test_gitignore_respected_without_git_dir(self, tmp_path):
+ (tmp_path / ".gitignore").write_text("ignored/\n*.log\n")
+ (tmp_path / "keep.py").write_text("print('hi')\n")
+ (tmp_path / "debug.log").write_text("noise\n")
+ (tmp_path / "ignored").mkdir()
+ (tmp_path / "ignored" / "junk.py").write_text("x = 1\n")
+
+ results = scan_files(target=tmp_path) # ← (F)
+ names = {r[0].name for r in results}
+ assert "keep.py" in names
+ assert "debug.log" not in names # ← (G)
+ assert "junk.py" not in names # ← (H)
| Mark | Annotation | Severity |
|---|---|---|
| F | Test exercises the regression directly: tmp_path is system tmp on Linux/macOS and typically has no ancestor .git. Fragility note in F-1 below. |
— |
| G | Covers the file-glob case (*.log). |
[POSITIVE] |
| H | Covers the directory case (ignored/). Both major gitignore pattern shapes get coverage. |
[POSITIVE] |
The test asserts the positive (keep.py in) and the negatives (debug.log, junk.py not in). That's the right shape — a test that only asserted exclusions could pass with scan_files returning [].
Findings
[MEDIUM] F-1 — Test is silently dependent on tmp_path having no ancestor .git
find_git_root walks upward from target.resolve() until it hits the filesystem root, looking for any .git. If a developer's test environment puts tmp_path somewhere under a git checkout (rare but possible — custom tmpdir configs, certain CI sandboxes, network-mounted home dirs with stray .git symlinks), the test would skip the new code path entirely. It would still likely pass because of how the patterns happen to be structured, but it would no longer be testing what its name claims.
Suggested fix: Either (a) explicitly verify no ancestor has .git at the start of the test, or (b) make the intent unambiguous by also asserting via a second call with no_gitignore=True that the included set differs:
results_no_ignore = scan_files(target=tmp_path, no_gitignore=True)
no_ignore_names = {r[0].name for r in results_no_ignore}
assert "debug.log" in no_ignore_names # confirms gitignore did the filtering
That second assertion locks the contract: "filtering happened because of gitignore handling," not "filtering happened, somehow."
[MEDIUM] F-2 — Behavior of fallback should be documented
The docstring of scan_files doesn't mention gitignore handling at all today. With this change, the rule "we'll honor a .gitignore in target even without .git" is now part of the contract. One line is enough:
"""Scan a directory for files, returning (path, mime_type, file_size) tuples.
…
When no_gitignore is False (default), .gitignore files are honored. The
gitignore root is the nearest ancestor containing .git, or the target
itself if no such ancestor exists. .gitignore files above the gitignore
root are not consulted.
"""
[LOW] F-3 — Unused parameter in load_gitignore_specs
def load_gitignore_specs(*, git_root: Path, target: Path) -> pathspec.PathSpec | None:
patterns: list[str] = []
target_resolved = target.resolve() # ← never read
...
target is no longer used inside this function. Pre-existing, not introduced by this PR — but the PR is the natural pass-by for it. Either remove the parameter and the dead line, or use it (e.g., to skip .gitignore files outside the target subtree if you wanted to make scoping stricter — though I'd argue you don't, because git itself doesn't).
If removing: also rename git_root → root while you're there, since the parameter is no longer git-specific in concept.
[LOW] F-4 — Style: git_root or target over the ternary
gitignore_root = git_root if git_root else target
# vs
gitignore_root = git_root or target
Path instances are always truthy, and find_git_root returns None or a Path, so the short form is both safe and idiomatic. Pure preference.
[LOW] F-5 — find_git_root accepts .git as either file or directory; commit message says "directory"
find_git_root checks (current / ".git").exists(), which is true for both directories and files. Worktrees use a .git file (containing gitdir: ...). The commit message says "no .git directory," but the code does the right thing for worktrees too. This is a wording nit on the commit message, not a code issue. The change correctly does not regress worktree handling.
[POSITIVE] F-6 — Minimal diff
Four-line change in production code, two of them pure renames, plus a focused test. Doesn't touch unrelated logic. Doesn't introduce new abstractions. The kind of fix that ages well.
[POSITIVE] F-7 — Test assertions cover both pattern shapes
*.log (file glob) and ignored/ (directory) are the two pattern syntaxes most users care about, and both are exercised. A .gitignore parser regression in either would fail this test.
Suggested commit-message tweak (optional)
Respect .gitignore when target has no .git directory or worktree marker
Tiny edit; keeps the message accurate for the worktree-file case the code already handles.
Pre-merge checklist
- [ ] Add the docstring note from F-2 (10 seconds).
- [ ] Optional: tighten the test per F-1 (one extra
no_gitignore=Truecall). - [ ] Optional: address F-3 in a follow-up cleanup commit.
- [ ] No security implications. No performance regression: when
git_rootisNoneand target has no.gitignore,load_gitignore_specswalks the target tree once and returnsNone. That walk is bounded by the same tree the mainos.walktraverses anyway, so worst case is one extra traversal of a directory the user already chose to scan.
Ready to merge after F-2.