# Is HTML "Strictly Better" Than Markdown for Claude Code?

By Corey J. Gallon · 2026-05-11 · https://gallon.me/is-html-really-strictly-better-than-markdown-for-claude-code-i-ran-the-numbers.html

---

Thariq Shihipar of Claude Code team fame posted an article last week, *"The Unreasonable Effectiveness of HTML"*, arguing that Claude should output HTML files by default for basically everything. PR reviews, postmortems, status reports, the lot. It's a good piece, racked up 8.3M views, and most of it I agree with. But in the replies, when somebody pushed back to say markdown's fine for text-heavy stuff, Thariq held the line: *"idk I think HTML is strictly better for all of that too."*
So I ran it all. Three of his use cases, instrumented end to end on Opus 4.7. What follows is what I measured, what surprised me, and where I ended up agreeing and disagreeing. Short version: he's right about half of it, and that half is *very* right. The other half I don't think holds up, and the data here is why.

## TL;DR

Thariq's [piece](https://x.com/trq212/status/2052809885763747935) is good. Read it before you read this. 

- **Visual tasks** (design exploration, prototypes, drag-and-drop widgets): HTML wins outright. Markdown can describe a design; HTML can show one. The 2-4x token cost is the price of admission, not a tradeoff.
- **Text-with-structure tasks** (PR reviews, technical explainers): both formats produce a real artifact. HTML adds polish for about 1.4-1.5x the dollar cost. In one of my two text cases (the rate-limiter explainer) the *markdown* surfaced more gotchas (12 vs 8) at lower cost. Which cuts against "strictly better."
- **Does Claude read HTML better than markdown?** (the question Thariq's argument quietly assumes a HTML-favorable answer to): two blinded judges scored Claude 7/7 on factual accuracy from both formats, identically. On explanatory specificity HTML got a ~7% edge, at 33% higher cost per call. Pick your poison.
- **The "1M context handles it" hand-wave:** at N=50 artifacts re-ingested K=3 times with cache evictions, HTML's overhead eats roughly the whole 1M window. Not nothing, but depending on the work being done maybe worth it.

Sample is small (N=1 per condition for content quality, three instrumented cases) and the methodology I used actually biases *against* markdown. Even with the thumb on the scale, "strictly better" isn't what I saw.

## What Thariq actually said

He walks through five buckets of tasks where he reckons HTML should be the default. Roughly:

1. Specs, planning, exploration (e.g. design exploration)
2. Code review & understanding (e.g. PR review with annotated diff)
3. Design & prototypes (e.g. an interactive checkout button)
4. Reports, research, learning (e.g. a rate limiter explainer)
5. Custom editing interfaces (e.g. drag-and-drop Linear ticket reorderer)

He flags three costs honestly: more tokens, slower generation, noisier diffs. He concludes the advantages outweigh them. Reasonable people can disagree on the conclusion.

What I want to look at is the reply to the comment. *"Strictly better"* is a specific, testable claim. That's the one I want to examine.

## There are actually two questions here

The article treats format choice as a knob: pay a bit extra in tokens, get nicer output, decide if it's worth it. That framing works for half the use cases and quietly stops working for the other half. Here's the split.

**Sometimes HTML can do something markdown structurally can't.** Render a six-mockup grid of UI designs. Run a slider that updates a preview in real time. Drag a card across a kanban. Markdown isn't a worse version of these; there is no markdown version. Asking "is the extra cost worth it?" is like asking whether the airfare to Mars is good value when the alternative is staying home and looking at a globe.

**Other times both formats can do the job, and HTML just costs more.** PR reviews. Postmortems. Technical explainers. Implementation plans. Markdown handles all of them. HTML renders them more attractively. The substance lives in either.

The "use HTML for everything" framing treats both questions like the second one. The case for HTML gets stronger, not weaker, when you split them apart.

## How I ran this

I picked three of Thariq's five use cases and instrumented them properly. Same prompt, both formats, capture every token. The other two are the categorical ones from the previous section; there's nothing fair to compare them against, so I just demo them at the end.

The three I instrumented:
1. **Design exploration**: onboarding screen, six approaches in a grid (his specs/planning bucket)
2. **PR review** of a real commit: `toks` commit `c6d70f9`, "Respect .gitignore when target has no .git directory"
3. **Rate limiter explainer** over the `slowapi` source: about 12K tokens of Python as substrate

The two I just demo:
4. Checkout button prototype (interactive sliders, animation, copy-as-prompt)
5. Linear ticket reorderer (drag-and-drop kanban with copy-as-markdown export)

For each instrumented case I ran three probes: generate the artifact, feed it back into a fresh agent session as context for a downstream task, and apply a semantic edit while capturing the diff. The first probe is the obvious one. The other two matter because that's where the cumulative token costs actually bite. Re-ingestion and editing happen *constantly* in real Claude Code workflows.

On the design exploration case I ran two methodologies side-by-side: Thariq's prompt unchanged versus a neutrally-phrased version of the same task. Why? Because his prompt literally says "create an HTML file." If you only do a "HTML" → "markdown" word-swap, you're asking Claude to make a markdown file in the shape of an HTML file, which is not the same as asking for markdown. Tracking? The difference between the two methodologies turned out to be larger than I expected: a 3.6x cost ratio under ablation, 2.1x under neutral phrasing. Worth being honest about.

> **What "artifact tokens" means below.** I ran `toks <file> --for claude` on each generated artifact, using Anthropic's own tokenizer to count the file the way Claude would see it if fed back as context. That number is different from `usage.output_tokens` (what the API reported as the generation length) and from `cache_creation_input_tokens` (artifact + system prompt + user message, all of it). The first matters for re-ingestion cost. The other two matter for other things.

**Where this is honestly a bit thin:**
- My reader-side reactions are N=1. One user (me), one session, I knew which file was which. No blinding, no randomization. Take them as gut reactions, not data.
- Generation runs are also N=1 per condition. Content-quality findings like "markdown surfaced more gotchas" might be variance, not signal. K=5 per condition would tell us. I didn't run K=5.
- Three instrumented cases is a sample, not a census.
- The agent-as-reader test below covers shared-content questions only, i.e. facts present in both artifacts. I didn't test whether HTML's verbose-priming or callout-salience changes output *quality* (beyond binary accuracy) in ways that matter downstream.
- I didn't test "colleagues actually read HTML more." That'd need users I don't have.
- **Methodology asymmetry** is load-bearing, and I'm flagging it here so the reader knows the thumb is on the scale before the data tables start. I ran two methodologies only on the design exploration case. The PR review and rate-limiter cases used Thariq's verbatim prompt structure, which is HTML-affording ("create an HTML artifact..."). That biases *against* the markdown artifact compared to a neutrally-phrased prompt. The cost ratios I report for those cases (1.51x for PR review, 1.44x for rate limiter) would likely shrink under a neutral prompt; how much, I didn't measure.

## Design exploration: this is where HTML earns its keep

**Thariq's prompt (verbatim):** *"I'm not sure what direction to take the onboarding screen. Generate 6 distinctly different approaches, varying layout, tone, and density, and lay them out as a single HTML file in a grid so I can compare them side by side. Label each with the tradeoff it's making."*

**Cost data (verbatim methodology):**

| Format | Artifact tokens | Output tokens | Generation time | Cost |
|---|---|---|---|---|
| HTML | 6,877 | 10,135 | 117 s | $0.51 |
| MD   | 1,896 | 3,284  | 49 s  | $0.27 |

**Cost data (format-neutral methodology):**

| Format | Artifact tokens | Output tokens | Generation time | Cost |
|---|---|---|---|---|
| HTML | 6,871 | 9,161 | 102 s | $0.48 |
| MD   | 3,332 | 4,769 | 73 s  | $0.34 |

Reading both artifacts side by side, my reaction was unambiguous:

> "HTML is **massively** nicer to read in these cases versus the markdown. It's much better ... it's not speed that makes the entire difference. It's quality and detail of visual output. This question is fundamentally visual and HTML allows for near-perfect visual communication of the ideas vs text-only in markdown."

The HTML rendered six high-fidelity onboarding mockups in iframes inside a responsive grid. The markdown produced a six-section text document that described what each mockup *would* look like.

<iframe srcdoc="<!doctype html>
<html lang=&quot;en&quot;>
<head>
<meta charset=&quot;utf-8&quot;>
<meta name=&quot;viewport&quot; content=&quot;width=device-width, initial-scale=1&quot;>
<title>Onboarding Screen — 6 Approaches</title>
<style>
  :root {
    --ink: #0f172a;
    --ink-2: #475569;
    --line: #e2e8f0;
    --bg: #f8fafc;
    --accent: #2563eb;
    --warn: #b45309;
  }
  * { box-sizing: border-box; }
  html, body { margin: 0; padding: 0; }
  body {
    font: 15px/1.5 -apple-system, BlinkMacSystemFont, &quot;Segoe UI&quot;, Roboto, &quot;Helvetica Neue&quot;, Arial, sans-serif;
    color: var(--ink);
    background: var(--bg);
    padding: 28px 24px 64px;
  }
  header.page {
    max-width: 1400px;
    margin: 0 auto 20px;
  }
  header.page h1 {
    margin: 0 0 6px;
    font-size: 22px;
    font-weight: 700;
    letter-spacing: -0.01em;
  }
  header.page p {
    margin: 0;
    color: var(--ink-2);
    font-size: 14px;
    max-width: 720px;
  }

  .grid {
    max-width: 1400px;
    margin: 0 auto;
    display: grid;
    grid-template-columns: repeat(3, 1fr);
    gap: 20px;
  }
  @media (max-width: 1100px) { .grid { grid-template-columns: repeat(2, 1fr); } }
  @media (max-width: 700px)  { .grid { grid-template-columns: 1fr; } }

  .card {
    background: #fff;
    border: 1px solid var(--line);
    border-radius: 12px;
    overflow: hidden;
    display: flex;
    flex-direction: column;
    box-shadow: 0 1px 2px rgba(15,23,42,0.04);
  }
  .card-label {
    padding: 12px 14px 10px;
    border-bottom: 1px solid var(--line);
    background: #fff;
  }
  .card-label .row {
    display: flex;
    align-items: baseline;
    justify-content: space-between;
    gap: 10px;
    margin-bottom: 4px;
  }
  .card-label .num {
    font-size: 11px;
    font-weight: 700;
    letter-spacing: 0.08em;
    color: var(--ink-2);
    text-transform: uppercase;
  }
  .card-label .name {
    font-size: 15px;
    font-weight: 600;
    color: var(--ink);
  }
  .card-label .tradeoff {
    font-size: 12.5px;
    color: var(--ink-2);
    margin: 2px 0 0;
  }
  .card-label .tradeoff b {
    color: var(--ink);
    font-weight: 600;
  }
  .frame-wrap {
    background: #f1f5f9;
    aspect-ratio: 4 / 3;
    position: relative;
  }
  .frame-wrap iframe {
    position: absolute;
    inset: 0;
    width: 100%;
    height: 100%;
    border: 0;
    background: #fff;
  }
</style>
</head>
<body>

<header class=&quot;page&quot;>
  <h1>Onboarding Screen — 6 Approaches</h1>
  <p>Each tile is a self-contained mockup of the same product moment (first launch, account just created). They vary in layout, tone, and information density. The label on each names the tradeoff that approach is choosing.</p>
</header>

<div class=&quot;grid&quot;>

  <!-- ============================================================ -->
  <!-- 1. MINIMAL: single field, single ask                          -->
  <!-- ============================================================ -->
  <article class=&quot;card&quot;>
    <div class=&quot;card-label&quot;>
      <div class=&quot;row&quot;>
        <span class=&quot;num&quot;>01</span>
        <span class=&quot;name&quot;>Minimal — single field</span>
      </div>
      <p class=&quot;tradeoff&quot;><b>Buys:</b> near-zero friction, fast to value. <b>Costs:</b> no signal of product depth, no personalization data captured.</p>
    </div>
    <div class=&quot;frame-wrap&quot;>
      <iframe title=&quot;Minimal&quot; srcdoc='<!doctype html><html><head><meta charset=utf-8><style>
        html,body{margin:0;height:100%;font:15px/1.5 -apple-system,BlinkMacSystemFont,&amp;quot;Segoe UI&amp;quot;,Roboto,sans-serif;color:#0f172a}
        .stage{height:100%;display:flex;align-items:center;justify-content:center;background:#fff;padding:24px}
        .pane{width:100%;max-width:360px;text-align:center}
        .logo{width:36px;height:36px;border-radius:9px;background:#0f172a;margin:0 auto 28px}
        h1{font-size:22px;font-weight:600;margin:0 0 8px;letter-spacing:-0.01em}
        p{color:#64748b;font-size:14px;margin:0 0 28px}
        .field{display:block;width:100%;padding:14px 16px;border:1px solid #e2e8f0;border-radius:10px;font-size:15px;outline:none}
        .field:focus{border-color:#2563eb}
        .hint{font-size:12px;color:#94a3b8;margin-top:14px}
      </style></head><body><div class=stage><div class=pane>
        <div class=logo></div>
        <h1>What should we call you?</h1>
        <p>One question. We will set up the rest as you go.</p>
        <input class=field placeholder=&quot;Your first name&quot;>
        <div class=hint>Press Enter to continue</div>
      </div></div></body></html>'></iframe>
    </div>
  </article>

  <!-- ============================================================ -->
  <!-- 2. WIZARD: multi-step, structured                             -->
  <!-- ============================================================ -->
  <article class=&quot;card&quot;>
    <div class=&quot;card-label&quot;>
      <div class=&quot;row&quot;>
        <span class=&quot;num&quot;>02</span>
        <span class=&quot;name&quot;>Wizard. Multi-step</span>
      </div>
      <p class=&quot;tradeoff&quot;><b>Buys:</b> thorough personalization, sets expectations. <b>Costs:</b> highest abandonment risk; each step is an exit ramp.</p>
    </div>
    <div class=&quot;frame-wrap&quot;>
      <iframe title=&quot;Wizard&quot; srcdoc='<!doctype html><html><head><meta charset=utf-8><style>
        html,body{margin:0;height:100%;font:14px/1.5 -apple-system,BlinkMacSystemFont,&amp;quot;Segoe UI&amp;quot;,Roboto,sans-serif;color:#0f172a;background:#fff}
        .wrap{height:100%;display:flex;flex-direction:column;padding:18px 22px}
        .top{display:flex;justify-content:space-between;align-items:center;font-size:12px;color:#64748b}
        .steps{display:flex;gap:6px;margin:10px 0 18px}
        .seg{flex:1;height:4px;background:#e2e8f0;border-radius:2px}
        .seg.done{background:#0f172a}
        .seg.now{background:#2563eb}
        h2{font-size:18px;margin:4px 0 4px;font-weight:600}
        .sub{color:#64748b;margin:0 0 14px;font-size:13px}
        .opts{display:grid;grid-template-columns:1fr 1fr;gap:8px}
        .opt{border:1px solid #e2e8f0;border-radius:8px;padding:10px 12px;font-size:13px;cursor:pointer}
        .opt.sel{border-color:#2563eb;background:#eff6ff}
        .footer{margin-top:auto;display:flex;justify-content:space-between;align-items:center;padding-top:14px}
        .back{font-size:13px;color:#64748b;background:none;border:0;cursor:pointer}
        .next{background:#0f172a;color:#fff;border:0;padding:9px 16px;border-radius:8px;font-size:13px;font-weight:500;cursor:pointer}
      </style></head><body><div class=wrap>
        <div class=top><span>Setup</span><span>Step 2 of 4</span></div>
        <div class=steps>
          <div class=&quot;seg done&quot;></div>
          <div class=&quot;seg now&quot;></div>
          <div class=seg></div>
          <div class=seg></div>
        </div>
        <h2>What kind of work do you do?</h2>
        <p class=sub>We will tailor templates and integrations to match.</p>
        <div class=opts>
          <div class=&quot;opt sel&quot;>Engineering</div>
          <div class=opt>Design</div>
          <div class=opt>Product</div>
          <div class=opt>Marketing</div>
          <div class=opt>Operations</div>
          <div class=opt>Something else</div>
        </div>
        <div class=footer>
          <button class=back>&amp;larr; Back</button>
          <button class=next>Continue</button>
        </div>
      </div></body></html>'></iframe>
    </div>
  </article>

  <!-- ============================================================ -->
  <!-- 3. SKIP-FIRST: get to value, defer setup                      -->
  <!-- ============================================================ -->
  <article class=&quot;card&quot;>
    <div class=&quot;card-label&quot;>
      <div class=&quot;row&quot;>
        <span class=&quot;num&quot;>03</span>
        <span class=&quot;name&quot;>Skip-first. Load sample data</span>
      </div>
      <p class=&quot;tradeoff&quot;><b>Buys:</b> fastest path to seeing the product work. <b>Costs:</b> users may never come back to configure; sample data can mislead.</p>
    </div>
    <div class=&quot;frame-wrap&quot;>
      <iframe title=&quot;Skip-first&quot; srcdoc='<!doctype html><html><head><meta charset=utf-8><style>
        html,body{margin:0;height:100%;font:14px/1.5 -apple-system,BlinkMacSystemFont,&amp;quot;Segoe UI&amp;quot;,Roboto,sans-serif;color:#0f172a;background:#fafafa}
        .wrap{height:100%;display:flex;align-items:center;justify-content:center;padding:24px}
        .card{background:#fff;border:1px solid #e2e8f0;border-radius:12px;padding:28px;max-width:380px;width:100%;text-align:center}
        h1{margin:0 0 8px;font-size:20px;font-weight:600}
        p{color:#64748b;margin:0 0 22px;font-size:13.5px}
        .primary{display:block;width:100%;background:#0f172a;color:#fff;border:0;padding:13px;border-radius:10px;font-size:14px;font-weight:600;cursor:pointer;margin-bottom:10px}
        .ghost{display:block;width:100%;background:#fff;color:#0f172a;border:1px solid #e2e8f0;padding:13px;border-radius:10px;font-size:14px;cursor:pointer}
        .or{font-size:12px;color:#94a3b8;margin:14px 0;letter-spacing:0.08em}
        .skip{margin-top:18px;font-size:13px;color:#2563eb;text-decoration:none;display:inline-block;border-bottom:1px solid #bfdbfe}
      </style></head><body><div class=wrap><div class=card>
        <h1>You are in.</h1>
        <p>Try it now with a sample workspace, or import your own data.</p>
        <button class=primary>Try with sample data</button>
        <button class=ghost>Import from CSV</button>
        <div class=or>OR</div>
        <a class=skip href=&quot;#&quot;>Skip and start from scratch &amp;rarr;</a>
      </div></div></body></html>'></iframe>
    </div>
  </article>

  <!-- ============================================================ -->
  <!-- 4. CONVERSATIONAL: chat intake                                -->
  <!-- ============================================================ -->
  <article class=&quot;card&quot;>
    <div class=&quot;card-label&quot;>
      <div class=&quot;row&quot;>
        <span class=&quot;num&quot;>04</span>
        <span class=&quot;name&quot;>Conversational. Chat intake</span>
      </div>
      <p class=&quot;tradeoff&quot;><b>Buys:</b> warm, human tone; lowers stakes of each ask. <b>Costs:</b> slow, sequential, hard to scan or skip ahead.</p>
    </div>
    <div class=&quot;frame-wrap&quot;>
      <iframe title=&quot;Conversational&quot; srcdoc='<!doctype html><html><head><meta charset=utf-8><style>
        html,body{margin:0;height:100%;font:14px/1.5 -apple-system,BlinkMacSystemFont,&amp;quot;Segoe UI&amp;quot;,Roboto,sans-serif;color:#0f172a;background:#f8fafc}
        .wrap{height:100%;display:flex;flex-direction:column;padding:16px 18px}
        .head{display:flex;align-items:center;gap:10px;margin-bottom:14px}
        .av{width:30px;height:30px;border-radius:50%;background:linear-gradient(135deg,#a78bfa,#60a5fa)}
        .who{font-size:13px;font-weight:600}
        .who small{display:block;color:#64748b;font-weight:400;font-size:11px}
        .stream{flex:1;display:flex;flex-direction:column;gap:8px;overflow:hidden}
        .msg{max-width:78%;padding:9px 12px;border-radius:14px;font-size:13px}
        .bot{background:#fff;border:1px solid #e2e8f0;align-self:flex-start;border-bottom-left-radius:4px}
        .me{background:#0f172a;color:#fff;align-self:flex-end;border-bottom-right-radius:4px}
        .chips{display:flex;flex-wrap:wrap;gap:6px;margin-top:4px}
        .chip{background:#fff;border:1px solid #cbd5e1;border-radius:14px;padding:5px 11px;font-size:12px;cursor:pointer}
        .input{margin-top:12px;display:flex;gap:8px;background:#fff;border:1px solid #e2e8f0;border-radius:22px;padding:6px 6px 6px 14px}
        .input input{flex:1;border:0;outline:none;font-size:13px;background:transparent}
        .send{background:#0f172a;color:#fff;border:0;border-radius:50%;width:32px;height:32px;cursor:pointer}
      </style></head><body><div class=wrap>
        <div class=head>
          <div class=av></div>
          <div class=who>Nova<small>Setup assistant</small></div>
        </div>
        <div class=stream>
          <div class=&quot;msg bot&quot;>Hey! Glad you are here. Mind if I ask a couple things to set up your space?</div>
          <div class=&quot;msg me&quot;>Sure</div>
          <div class=&quot;msg bot&quot;>What should I call you?</div>
          <div class=&quot;msg me&quot;>Sam</div>
          <div class=&quot;msg bot&quot;>Nice to meet you, Sam. Is this for personal use or for a team?</div>
          <div class=chips>
            <span class=chip>Just me</span>
            <span class=chip>A small team</span>
            <span class=chip>A whole company</span>
          </div>
        </div>
        <div class=input>
          <input placeholder=&quot;Type a reply...&quot;>
          <button class=send>&amp;uarr;</button>
        </div>
      </div></body></html>'></iframe>
    </div>
  </article>

  <!-- ============================================================ -->
  <!-- 5. DASHBOARD-PREVIEW: app behind, checklist overlay           -->
  <!-- ============================================================ -->
  <article class=&quot;card&quot;>
    <div class=&quot;card-label&quot;>
      <div class=&quot;row&quot;>
        <span class=&quot;num&quot;>05</span>
        <span class=&quot;name&quot;>Dashboard-preview. Checklist overlay</span>
      </div>
      <p class=&quot;tradeoff&quot;><b>Buys:</b> shows the destination immediately; setup feels like progress, not a gate. <b>Costs:</b> dense and busy; can overwhelm on first impression.</p>
    </div>
    <div class=&quot;frame-wrap&quot;>
      <iframe title=&quot;Dashboard-preview&quot; srcdoc='<!doctype html><html><head><meta charset=utf-8><style>
        html,body{margin:0;height:100%;font:13px/1.4 -apple-system,BlinkMacSystemFont,&amp;quot;Segoe UI&amp;quot;,Roboto,sans-serif;color:#0f172a;background:#f1f5f9;overflow:hidden}
        .app{height:100%;display:grid;grid-template-columns:140px 1fr;position:relative}
        .side{background:#0f172a;color:#cbd5e1;padding:14px 12px;font-size:12px}
        .side .b{color:#fff;font-weight:600;margin-bottom:14px}
        .side .it{padding:6px 8px;border-radius:6px;margin-bottom:2px;cursor:pointer}
        .side .it.on{background:rgba(255,255,255,0.08);color:#fff}
        .main{padding:14px 16px;background:#fff}
        .main h2{margin:0 0 10px;font-size:16px}
        .stats{display:grid;grid-template-columns:1fr 1fr 1fr;gap:8px;margin-bottom:12px}
        .stat{background:#f8fafc;border:1px solid #e2e8f0;border-radius:8px;padding:10px}
        .stat b{display:block;font-size:18px}
        .stat span{color:#64748b;font-size:11px}
        .row{height:10px;background:#e2e8f0;border-radius:3px;margin-bottom:6px}
        .row.s{width:60%}
        .row.m{width:80%}
        .overlay{position:absolute;right:14px;bottom:14px;width:240px;background:#fff;border:1px solid #e2e8f0;border-radius:12px;box-shadow:0 12px 30px rgba(15,23,42,0.18);padding:14px}
        .overlay h3{margin:0 0 4px;font-size:13px}
        .overlay .p{font-size:11px;color:#64748b;margin:0 0 10px}
        .bar{height:5px;background:#e2e8f0;border-radius:3px;overflow:hidden;margin-bottom:10px}
        .bar i{display:block;width:25%;height:100%;background:#22c55e}
        .ck{display:flex;align-items:center;gap:8px;font-size:12px;padding:4px 0}
        .box{width:14px;height:14px;border:1.5px solid #cbd5e1;border-radius:4px;flex-shrink:0}
        .box.done{background:#22c55e;border-color:#22c55e}
        .ck.done{color:#94a3b8;text-decoration:line-through}
      </style></head><body><div class=app>
        <div class=side>
          <div class=b>Acme</div>
          <div class=&quot;it on&quot;>Dashboard</div>
          <div class=it>Projects</div>
          <div class=it>Inbox</div>
          <div class=it>Reports</div>
          <div class=it>Settings</div>
        </div>
        <div class=main>
          <h2>Welcome, Sam</h2>
          <div class=stats>
            <div class=stat><b>0</b><span>Projects</span></div>
            <div class=stat><b>0</b><span>Tasks</span></div>
            <div class=stat><b>1</b><span>Members</span></div>
          </div>
          <div class=&quot;row m&quot;></div>
          <div class=&quot;row s&quot;></div>
          <div class=&quot;row m&quot;></div>
          <div class=&quot;row s&quot;></div>
        </div>
        <div class=overlay>
          <h3>Get set up (1 of 4)</h3>
          <p class=p>Two minutes. You can finish later.</p>
          <div class=bar><i></i></div>
          <div class=&quot;ck done&quot;><span class=&quot;box done&quot;></span>Create account</div>
          <div class=ck><span class=box></span>Invite a teammate</div>
          <div class=ck><span class=box></span>Connect a tool</div>
          <div class=ck><span class=box></span>Start a project</div>
        </div>
      </div></body></html>'></iframe>
    </div>
  </article>

  <!-- ============================================================ -->
  <!-- 6. PERSONA PICKER: pick a role, get a preset                  -->
  <!-- ============================================================ -->
  <article class=&quot;card&quot;>
    <div class=&quot;card-label&quot;>
      <div class=&quot;row&quot;>
        <span class=&quot;num&quot;>06</span>
        <span class=&quot;name&quot;>Persona picker. Pick a role</span>
      </div>
      <p class=&quot;tradeoff&quot;><b>Buys:</b> instant personalization from one tap. <b>Costs:</b> forces a category choice users may not have ready; hard to fit edge cases.</p>
    </div>
    <div class=&quot;frame-wrap&quot;>
      <iframe title=&quot;Persona&quot; srcdoc='<!doctype html><html><head><meta charset=utf-8><style>
        html,body{margin:0;height:100%;font:14px/1.5 -apple-system,BlinkMacSystemFont,&amp;quot;Segoe UI&amp;quot;,Roboto,sans-serif;color:#0f172a;background:#fff}
        .wrap{height:100%;display:flex;flex-direction:column;padding:20px 22px}
        h1{font-size:22px;margin:0 0 4px;font-weight:700;letter-spacing:-0.01em}
        p.sub{color:#64748b;margin:0 0 16px;font-size:13px}
        .grid{display:grid;grid-template-columns:1fr 1fr;gap:10px;flex:1}
        .tile{border:1.5px solid #e2e8f0;border-radius:12px;padding:12px;display:flex;flex-direction:column;justify-content:space-between;cursor:pointer;transition:border-color .15s}
        .tile:hover{border-color:#94a3b8}
        .tile.sel{border-color:#0f172a;background:#0f172a;color:#fff}
        .tile .ico{width:28px;height:28px;border-radius:8px;background:#f1f5f9;display:flex;align-items:center;justify-content:center;font-size:14px}
        .tile.sel .ico{background:rgba(255,255,255,0.12)}
        .tile .name{font-weight:600;font-size:14px;margin-top:18px}
        .tile .desc{font-size:11.5px;opacity:0.75;margin-top:2px;line-height:1.35}
        .foot{display:flex;justify-content:space-between;align-items:center;margin-top:14px}
        .foot a{font-size:12.5px;color:#64748b;text-decoration:none}
        .go{background:#0f172a;color:#fff;border:0;padding:9px 18px;border-radius:8px;font-size:13px;font-weight:600;cursor:pointer}
      </style></head><body><div class=wrap>
        <h1>Who are you, mostly?</h1>
        <p class=sub>Pick the closest match. We will preset your workspace.</p>
        <div class=grid>
          <div class=tile><div class=ico>&amp;#9881;</div><div><div class=name>Builder</div><div class=desc>Code, tickets, deploys.</div></div></div>
          <div class=&quot;tile sel&quot;><div class=ico>&amp;#9998;</div><div><div class=name>Maker</div><div class=desc>Drafts, files, calendars.</div></div></div>
          <div class=tile><div class=ico>&amp;#9742;</div><div><div class=name>Connector</div><div class=desc>People, threads, follow-ups.</div></div></div>
          <div class=tile><div class=ico>&amp;#9776;</div><div><div class=name>Operator</div><div class=desc>Lists, runbooks, status.</div></div></div>
        </div>
        <div class=foot>
          <a href=&quot;#&quot;>Not sure &amp;mdash; show me everything</a>
          <button class=go>Use this preset &amp;rarr;</button>
        </div>
      </div></body></html>'></iframe>
    </div>
  </article>

</div>
</body>
</html>
" height="700" width="100%" style="border:1px solid #d0d7de;border-radius:6px;margin:1.5em 0"></iframe>

*The HTML artifact, rendered above. Six approaches, side by side. This is the thing markdown structurally can't do.*

**Three things worth noting from this case:**

- *Markdown leaks HTML, but only when the prompt corners it.* When I ran Thariq's verbatim prompt (with its "single markdown file in a grid" phrasing), the model embedded 10 HTML tags into the "markdown" output (`<table>`, `<tr>`, `<td>`). Of course it did; there's no markdown way to lay out a grid. When I dropped "in a grid" from the prompt, the markdown came out clean. Zero HTML tags. So if you've ever wondered why your markdown agents sometimes spit out HTML, here's one answer: you asked them to do something markdown can't do.
- *Methodology shifts the cost ratio.* Verbatim ablation: 3.63x artifact tokens; format-neutral: 2.06x. Single-number ratios reported without methodology disclosure should be read skeptically.
- *Generation time partially supports Thariq's "2-4x slower" admission.* Measured 1.4-2.4x depending on methodology. The high end of his range is reachable, the low end is below what I measured.

The call: HTML wins outright. The token cost isn't a tradeoff, it's just what the capability costs.

## PR review: a fair fight, and HTML costs about 40% more for the polish

> **Methodology note:** this case used verbatim ablation only (the format-neutral methodology was only run on the design exploration). The verbatim prompt structure ("create an HTML artifact...") is HTML-affording and biases against markdown. The 1.51x cost ratio reported here would likely shrink under a neutrally-phrased prompt; I didn't measure by how much.

**Thariq's prompt (lightly adapted for my substrate):** *"Help me review this PR by creating an HTML artifact that describes it. I'm not familiar with the gitignore/path-resolution logic so brace on that. Render the actual diff with inline margin annotations, color-code findings by severity..."*

**Substrate:** `toks` commit c6d70f9, a real bug fix (~30 lines of code change plus a new test).

**Cost data:**

| Format | Artifact tokens | Output tokens | Generation time | Cost |
|---|---|---|---|---|
| HTML | 11,223 | 19,308 | 234 s | $0.82 |
| MD   | 3,810  | 10,437 | 154 s | $0.54 |

Ratios: 2.95x artifact tokens, 1.85x output tokens, 1.52x time, 1.51x cost.

The markdown PR review is, frankly, a real PR review. TL;DR verdict at the top ("Approve with two non-blocking notes"), a severity legend in a markdown table, a structured walkthrough of the gitignore logic, line-by-line diff annotations, a test-coverage assessment. If a teammate sent me this in a Slack DM I'd be perfectly happy.

HTML adds a color-coded verdict bar, rendered diffs with green/red highlighting and line numbers, eight severity-tagged finding cards (Pass / Concern / Nit), and inline severity callouts on the diff. Prettier. More navigable. The substance is the same.

<iframe srcdoc="<!DOCTYPE html>
<html lang=&quot;en&quot;>
<head>
<meta charset=&quot;UTF-8&quot;>
<title>PR Review &amp;mdash; toks c6d70f9: Respect .gitignore when target has no .git directory</title>
<style>
  :root {
    --bg: #ffffff;
    --fg: #1f2328;
    --muted: #59636e;
    --border: #d0d7de;
    --code-bg: #f6f8fa;
    --diff-add-bg: #dafbe1;
    --diff-add-bar: #2da44e;
    --diff-del-bg: #ffebe9;
    --diff-del-bar: #cf222e;
    --diff-ctx-bg: #ffffff;

    --sev-pass: #2da44e;
    --sev-pass-bg: #dcfce7;
    --sev-nit:  #0969da;
    --sev-nit-bg:  #ddf4ff;
    --sev-warn: #bf8700;
    --sev-warn-bg: #fff8c5;
    --sev-block:#cf222e;
    --sev-block-bg:#ffebe9;
  }
  html { font-family: -apple-system, BlinkMacSystemFont, &quot;Segoe UI&quot;, Helvetica, Arial, sans-serif; color: var(--fg); background: var(--bg); }
  body { max-width: 1200px; margin: 0 auto; padding: 32px 24px 64px; line-height: 1.5; }
  h1 { font-size: 24px; margin: 0 0 4px 0; }
  h2 { font-size: 18px; margin: 36px 0 12px 0; padding-bottom: 6px; border-bottom: 1px solid var(--border); }
  h3 { font-size: 15px; margin: 18px 0 8px 0; }
  code, pre, .mono { font-family: ui-monospace, SFMono-Regular, &quot;SF Mono&quot;, Menlo, Consolas, monospace; font-size: 13px; }
  code { background: var(--code-bg); padding: 1px 5px; border-radius: 4px; }
  pre { background: var(--code-bg); padding: 12px 14px; border-radius: 6px; overflow-x: auto; margin: 8px 0; }
  .meta { color: var(--muted); font-size: 13px; margin-bottom: 18px; }
  .meta .sha { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; }

  /* Verdict bar */
  .verdict {
    display: flex; align-items: stretch; gap: 16px; margin-top: 8px;
    border: 1px solid var(--border); border-radius: 8px; overflow: hidden;
  }
  .verdict .stripe { width: 8px; background: var(--sev-pass); }
  .verdict .body { padding: 14px 16px; flex: 1; }
  .verdict .label { font-weight: 600; color: var(--sev-pass); font-size: 13px; letter-spacing: 0.04em; text-transform: uppercase; }
  .verdict .summary { margin-top: 4px; }

  /* Severity legend */
  .legend { display: flex; flex-wrap: wrap; gap: 8px; margin: 14px 0 8px 0; }
  .legend .pill { display: inline-flex; align-items: center; gap: 8px; font-size: 12px; padding: 4px 10px; border-radius: 999px; border: 1px solid var(--border); background: #fff; }
  .legend .dot { width: 10px; height: 10px; border-radius: 50%; display: inline-block; }
  .dot.pass  { background: var(--sev-pass); }
  .dot.nit   { background: var(--sev-nit); }
  .dot.warn  { background: var(--sev-warn); }
  .dot.block { background: var(--sev-block); }

  /* Brace box (background explainer) */
  .brace { border: 1px solid var(--border); border-left: 4px solid var(--sev-nit); background: #f8fbff; padding: 14px 18px; border-radius: 6px; margin: 12px 0; }
  .brace h3 { margin-top: 0; color: var(--sev-nit); }
  .brace p { margin: 8px 0; }
  .brace ul { margin: 6px 0 6px 0; padding-left: 22px; }
  .brace li { margin: 4px 0; }

  /* Filesystem diagram */
  .fsdiag { display: grid; grid-template-columns: 1fr 1fr; gap: 14px; margin: 12px 0; }
  .fsdiag .box { border: 1px solid var(--border); border-radius: 6px; padding: 12px 14px; background: #fff; }
  .fsdiag .box h4 { margin: 0 0 8px 0; font-size: 13px; color: var(--muted); text-transform: uppercase; letter-spacing: 0.04em; }
  .tree { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: 12.5px; line-height: 1.55; white-space: pre; }
  .tree .ignored { color: var(--sev-block); text-decoration: line-through; }
  .tree .scanned { color: var(--sev-pass); font-weight: 600; }
  .tree .root    { color: var(--sev-nit); font-weight: 600; }
  .tree .target  { background: #fff3b0; padding: 0 4px; }

  /* Annotated diff: 2-column grid (diff on left, callouts on right) */
  .diff-block { margin: 14px 0 24px 0; }
  .hunk-header {
    background: #ddf4ff; color: #0969da; padding: 6px 10px; font-family: ui-monospace, SFMono-Regular, Menlo, monospace;
    font-size: 12.5px; border-radius: 6px 6px 0 0; border: 1px solid var(--border); border-bottom: 0;
  }
  .annotated {
    display: grid; grid-template-columns: minmax(0, 1.55fr) minmax(0, 1fr); gap: 0;
    border: 1px solid var(--border); border-top: 0; border-radius: 0 0 6px 6px; overflow: hidden;
  }
  .annotated .diff { background: var(--code-bg); padding: 0; min-width: 0; }
  .annotated .notes { background: #fafbfc; padding: 0; border-left: 1px solid var(--border); display: flex; flex-direction: column; min-width: 0; }
  .row { display: grid; grid-template-columns: 36px 36px 1fr; align-items: stretch; }
  .row .gutter { background: #eaeef2; color: var(--muted); text-align: right; padding: 2px 6px; font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: 11.5px; user-select: none; }
  .row .sign { padding: 2px 6px; text-align: center; font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: 12.5px; }
  .row .code { padding: 2px 8px; white-space: pre; overflow-x: auto; font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: 12.5px; }
  .row.ctx { background: #ffffff; }
  .row.add { background: var(--diff-add-bg); }
  .row.add .sign { color: var(--diff-add-bar); }
  .row.del { background: var(--diff-del-bg); }
  .row.del .sign { color: var(--diff-del-bar); }
  .row.tag { background: #fff8c5; }
  .row.tag .gutter, .row.tag .sign { background: transparent; color: var(--sev-warn); font-weight: 600; }

  /* annotation cards stacked in right pane */
  .ann {
    border-left: 4px solid var(--sev-nit); background: #fff; padding: 10px 12px; margin: 0; flex: 0 0 auto;
    border-bottom: 1px solid var(--border);
  }
  .ann:last-child { border-bottom: 0; }
  .ann.pass  { border-left-color: var(--sev-pass);  background: var(--sev-pass-bg); }
  .ann.nit   { border-left-color: var(--sev-nit);   background: var(--sev-nit-bg); }
  .ann.warn  { border-left-color: var(--sev-warn);  background: var(--sev-warn-bg); }
  .ann.block { border-left-color: var(--sev-block); background: var(--sev-block-bg); }
  .ann .head { font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; font-weight: 700; margin-bottom: 4px; }
  .ann.pass  .head { color: var(--sev-pass); }
  .ann.nit   .head { color: var(--sev-nit); }
  .ann.warn  .head { color: var(--sev-warn); }
  .ann.block .head { color: var(--sev-block); }
  .ann .body { font-size: 13px; }
  .ann .body code { background: rgba(255,255,255,0.6); }
  .ann .anchor { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: 11px; color: var(--muted); margin-bottom: 2px; }

  /* Findings table */
  .findings { display: flex; flex-direction: column; gap: 10px; }
  .finding {
    display: grid; grid-template-columns: 8px 1fr; gap: 0; border: 1px solid var(--border); border-radius: 6px; overflow: hidden; background: #fff;
  }
  .finding .stripe { width: 8px; }
  .finding.pass  .stripe { background: var(--sev-pass); }
  .finding.nit   .stripe { background: var(--sev-nit); }
  .finding.warn  .stripe { background: var(--sev-warn); }
  .finding.block .stripe { background: var(--sev-block); }
  .finding .body { padding: 12px 14px; }
  .finding .head { display: flex; align-items: center; gap: 10px; margin-bottom: 4px; }
  .finding .badge { font-size: 11px; text-transform: uppercase; letter-spacing: 0.06em; font-weight: 700; padding: 2px 8px; border-radius: 999px; }
  .finding.pass  .badge { background: var(--sev-pass-bg);  color: var(--sev-pass); }
  .finding.nit   .badge { background: var(--sev-nit-bg);   color: var(--sev-nit); }
  .finding.warn  .badge { background: var(--sev-warn-bg);  color: var(--sev-warn); }
  .finding.block .badge { background: var(--sev-block-bg); color: var(--sev-block); }
  .finding .title { font-weight: 600; }
  .finding .where { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: 12px; color: var(--muted); }
  .finding p { margin: 6px 0; }

  /* Test coverage matrix */
  table.cov { border-collapse: collapse; width: 100%; margin: 8px 0; font-size: 13px; }
  table.cov th, table.cov td { border: 1px solid var(--border); padding: 8px 10px; text-align: left; vertical-align: top; }
  table.cov th { background: var(--code-bg); font-weight: 600; }
  .yes { color: var(--sev-pass); font-weight: 600; }
  .no  { color: var(--sev-warn); font-weight: 600; }

  .footnote { color: var(--muted); font-size: 12px; margin-top: 8px; }
</style>
</head>
<body>

<h1>PR Review &amp;mdash; Respect <code>.gitignore</code> when target has no <code>.git</code> directory</h1>
<div class=&quot;meta&quot;>
  Repository <code>toks</code> &amp;middot; commit <span class=&quot;sha&quot;>c6d70f9</span> &amp;middot; author Corey Gallon &amp;middot; 2026-04-29
</div>

<div class=&quot;verdict&quot;>
  <div class=&quot;stripe&quot;></div>
  <div class=&quot;body&quot;>
    <div class=&quot;label&quot;>Approve with notes</div>
    <div class=&quot;summary&quot;>
      The fix is small, correct, and targeted. It closes a real bug:
      <code>.gitignore</code> was previously a no-op for any directory that didn't sit under a <code>.git</code> root.
      The new test exercises the fixed path. Two pre-existing rough edges in the surrounding logic
      are surfaced below as concerns&amp;mdash;not blockers&amp;mdash;because the PR description's motivating example
      (<code>.venv</code> being scanned) intersects them and a future reader will want to know.
    </div>
  </div>
</div>

<div class=&quot;legend&quot;>
  <span class=&quot;pill&quot;><span class=&quot;dot pass&quot;></span>Pass &amp;mdash; nothing to change</span>
  <span class=&quot;pill&quot;><span class=&quot;dot nit&quot;></span>Nit &amp;mdash; style/clarity</span>
  <span class=&quot;pill&quot;><span class=&quot;dot warn&quot;></span>Concern &amp;mdash; worth discussing</span>
  <span class=&quot;pill&quot;><span class=&quot;dot block&quot;></span>Blocker &amp;mdash; must fix</span>
</div>

<h2>Background: how gitignore resolution actually works here</h2>

<div class=&quot;brace&quot;>
  <h3>Two functions, two directions of walk</h3>
  <p>The scanner relies on two helpers in <code>src/toks/scanner.py</code>. To read this PR you have to hold both in your head at once:</p>
  <ul>
    <li><b><code>find_git_root(start=target)</code></b> walks <b>upward</b> from the target, looking for a <code>.git</code> directory in each parent. Returns that parent, or <code>None</code> if it reaches the filesystem root without finding one.</li>
    <li><b><code>load_gitignore_specs(git_root, target)</code></b> walks <b>downward</b> from <code>git_root</code> via <code>os.walk(git_root)</code>, collecting every <code>.gitignore</code> it finds, and prefixing each pattern with the <code>.gitignore</code>'s relative directory so a nested <code>foo/.gitignore</code> rule like <code>build/</code> becomes <code>foo/build/</code>.</li>
  </ul>
  <p>In the scan loop, each candidate file is reduced to a path <b>relative to that same root</b> via <code>file_path.relative_to(git_root)</code> before being matched against the spec. <i>This is the critical invariant</i>: the root used to <i>build</i> the spec must be the same root used to <i>relativize</i> the candidate, or every match silently fails (or, worse, raises <code>ValueError</code> from <code>relative_to</code>).</p>
  <p>The bug being fixed: when <code>find_git_root</code> returned <code>None</code>, the old code skipped <code>load_gitignore_specs</code> entirely. So a project with a <code>.gitignore</code> but no <code>git init</code> got no filtering at all.</p>
</div>

<div class=&quot;fsdiag&quot;>
  <div class=&quot;box&quot;>
    <h4>Before this PR (no <code>.git</code> upstream)</h4>
    <pre class=&quot;tree&quot;>/projects/<span class=&quot;target&quot;>myproj</span>/        &amp;larr; target
  .gitignore         (says: .venv/)
  src/
    main.py          <span class=&quot;scanned&quot;>scanned</span>
  .venv/
    lib/site-pkgs/
      django/...     <span class=&quot;scanned&quot;>scanned (bug!)</span>
      numpy/...      <span class=&quot;scanned&quot;>scanned (bug!)</span></pre>
    <p class=&quot;footnote&quot;><code>find_git_root</code> returns <code>None</code> &amp;rarr; <code>gitignore_spec</code> stays <code>None</code> &amp;rarr; the <code>.gitignore</code> is silently ignored.</p>
  </div>
  <div class=&quot;box&quot;>
    <h4>After this PR (no <code>.git</code> upstream)</h4>
    <pre class=&quot;tree&quot;>/projects/<span class=&quot;target&quot;>myproj</span>/        &amp;larr; <span class=&quot;root&quot;>target = gitignore_root</span>
  .gitignore         (says: .venv/)
  src/
    main.py          <span class=&quot;scanned&quot;>scanned</span>
  .venv/
    lib/site-pkgs/
      django/...     <span class=&quot;ignored&quot;>ignored</span>
      numpy/...      <span class=&quot;ignored&quot;>ignored</span></pre>
    <p class=&quot;footnote&quot;>Fallback: <code>gitignore_root = target</code>. <code>load_gitignore_specs</code> walks the target subtree, builds a spec from the local <code>.gitignore</code>, and the per-file check now matches.</p>
  </div>
</div>

<h2>Annotated diff &amp;mdash; <code>src/toks/scanner.py</code></h2>

<div class=&quot;diff-block&quot;>
  <div class=&quot;hunk-header&quot;>@@ -104,11 +104,11 @@ def scan_files(...):</div>
  <div class=&quot;annotated&quot;>
    <div class=&quot;diff&quot;>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>104</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>    raise ValueError(f&quot;Not a directory: {target}&quot;)</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>105</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;> </div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>106</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>    gitignore_spec = None</div></div>
      <div class=&quot;row del&quot;><div class=&quot;gutter&quot;>107</div><div class=&quot;sign&quot;>-</div><div class=&quot;code&quot;>    git_root = None</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>107</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>    gitignore_root = None</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>108</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>    if not no_gitignore:</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>109</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>        git_root = find_git_root(start=target)</div></div>
      <div class=&quot;row del&quot;><div class=&quot;gutter&quot;>110</div><div class=&quot;sign&quot;>-</div><div class=&quot;code&quot;>        if git_root:</div></div>
      <div class=&quot;row del&quot;><div class=&quot;gutter&quot;>111</div><div class=&quot;sign&quot;>-</div><div class=&quot;code&quot;>            gitignore_spec = load_gitignore_specs(git_root=git_root, target=target)</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>110</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        gitignore_root = git_root if git_root else target</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>111</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        gitignore_spec = load_gitignore_specs(git_root=gitignore_root, target=target)</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>112</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;> </div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>113</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>    results: list[tuple[Path, str, int]] = []</div></div>
    </div>
    <div class=&quot;notes&quot;>
      <div class=&quot;ann pass&quot;>
        <div class=&quot;anchor&quot;>line 107 &amp;mdash; rename</div>
        <div class=&quot;head&quot;>Pass &amp;mdash; better name</div>
        <div class=&quot;body&quot;>
          Renaming <code>git_root</code> &amp;rarr; <code>gitignore_root</code> is the right call. After the fix the variable can hold a path that has nothing to do with git (it can just be the target). The new name reflects what it's <i>used for</i> rather than where it came from.
        </div>
      </div>
      <div class=&quot;ann pass&quot;>
        <div class=&quot;anchor&quot;>lines 110&amp;ndash;111 &amp;mdash; the fix</div>
        <div class=&quot;head&quot;>Pass &amp;mdash; correct fallback</div>
        <div class=&quot;body&quot;>
          <code>git_root if git_root else target</code> establishes a non-null root in every branch where <code>no_gitignore</code> is false. The same value is then threaded into <code>load_gitignore_specs</code> as <code>git_root=&amp;hellip;</code>. The keyword name in the callee now reads as a slight misnomer (it's not necessarily a git root anymore), but that's a follow-up rename, not a bug.
        </div>
      </div>
      <div class=&quot;ann warn&quot;>
        <div class=&quot;anchor&quot;>line 109 &amp;mdash; pre-existing</div>
        <div class=&quot;head&quot;>Concern &amp;mdash; upstream <code>.git</code> can hijack the root</div>
        <div class=&quot;body&quot;>
          <code>find_git_root</code> walks all the way up to <code>/</code>. If the user has any unrelated <code>.git</code> upstream of the target (e.g. a dotfiles repo at <code>~/.git</code>, or a parent monorepo), it becomes the root. <code>load_gitignore_specs</code> then <code>os.walk</code>s the entire ancestor tree to harvest every <code>.gitignore</code> under it. This was true before the PR too &amp;mdash; the PR doesn't make it worse &amp;mdash; but it's worth surfacing because the new fallback only kicks in when <i>no</i> <code>.git</code> is found, and the more common surprising case is finding the wrong one.
        </div>
      </div>
      <div class=&quot;ann nit&quot;>
        <div class=&quot;anchor&quot;>line 111 &amp;mdash; small style</div>
        <div class=&quot;head&quot;>Nit &amp;mdash; consider renaming the parameter too</div>
        <div class=&quot;body&quot;>
          <code>load_gitignore_specs(git_root=&amp;hellip;)</code> still takes a parameter named <code>git_root</code>. After this change, callers pass either a true git root or the target. Renaming the parameter to <code>root</code> would close the loop on the rename you started in this PR. Optional &amp;mdash; do it now or in a follow-up.
        </div>
      </div>
    </div>
  </div>
</div>

<div class=&quot;diff-block&quot;>
  <div class=&quot;hunk-header&quot;>@@ -128,8 +128,8 @@ def scan_files(...):</div>
  <div class=&quot;annotated&quot;>
    <div class=&quot;diff&quot;>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>128</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>            if file_path.is_symlink() and file_path.is_dir():</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>129</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>                continue</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>130</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;> </div></div>
      <div class=&quot;row del&quot;><div class=&quot;gutter&quot;>131</div><div class=&quot;sign&quot;>-</div><div class=&quot;code&quot;>            if gitignore_spec and git_root:</div></div>
      <div class=&quot;row del&quot;><div class=&quot;gutter&quot;>132</div><div class=&quot;sign&quot;>-</div><div class=&quot;code&quot;>                rel = file_path.relative_to(git_root)</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>131</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>            if gitignore_spec and gitignore_root:</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>132</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>                rel = file_path.relative_to(gitignore_root)</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>133</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>                if gitignore_spec.match_file(str(rel)):</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>134</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>                    continue</div></div>
    </div>
    <div class=&quot;notes&quot;>
      <div class=&quot;ann pass&quot;>
        <div class=&quot;anchor&quot;>line 132 &amp;mdash; invariant preserved</div>
        <div class=&quot;head&quot;>Pass &amp;mdash; root used to build = root used to relativize</div>
        <div class=&quot;body&quot;>
          The critical invariant from the brace above is intact: the spec is built with <code>gitignore_root</code> in <code>load_gitignore_specs</code>, and candidate files are relativized to the <i>same</i> <code>gitignore_root</code> here. If a future change splits these, that's where bugs would creep in.
        </div>
      </div>
      <div class=&quot;ann nit&quot;>
        <div class=&quot;anchor&quot;>line 131 &amp;mdash; defensive check</div>
        <div class=&quot;head&quot;>Nit &amp;mdash; <code>and gitignore_root</code> is now redundant</div>
        <div class=&quot;body&quot;>
          After the fix, whenever <code>gitignore_spec</code> is non-<code>None</code>, <code>gitignore_root</code> is also set (the only path that produces a spec assigns the root first). The <code>and gitignore_root</code> guard is harmless but no longer carries information. You could drop it, or keep it as defensive code &amp;mdash; not worth blocking on.
        </div>
      </div>
      <div class=&quot;ann warn&quot;>
        <div class=&quot;anchor&quot;>whole hunk &amp;mdash; pre-existing perf</div>
        <div class=&quot;head&quot;>Concern &amp;mdash; gitignore filtering is per-file, not per-directory</div>
        <div class=&quot;body&quot;>
          The walk in <code>scan_files</code> only consults the spec for <i>files</i>. Directory pruning is hard-coded to <code>.git</code> and symlinks. So even after this fix, if <code>.gitignore</code> contains <code>.venv/</code>, <code>os.walk</code> still descends into <code>.venv/</code>, stats every file, and then drops them one-by-one via <code>match_file</code>. The output is correct; the walk is slow on a directory full of <code>site-packages</code>. The PR description names <code>.venv</code> specifically, which is exactly the case where users will notice the cost. Worth a follow-up: prune <code>dirnames</code> against the spec before recursing.
        </div>
      </div>
    </div>
  </div>
</div>

<h2>Annotated diff &amp;mdash; <code>tests/test_scanner.py</code></h2>

<div class=&quot;diff-block&quot;>
  <div class=&quot;hunk-header&quot;>@@ -95,3 +95,16 @@ class TestScanFiles:</div>
  <div class=&quot;annotated&quot;>
    <div class=&quot;diff&quot;>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>95</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>        empty.mkdir()</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>96</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>        results = scan_files(target=empty)</div></div>
      <div class=&quot;row ctx&quot;><div class=&quot;gutter&quot;>97</div><div class=&quot;sign&quot;> </div><div class=&quot;code&quot;>        assert results == []</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>98</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;> </div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>99</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>    def test_gitignore_respected_without_git_dir(self, tmp_path):</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>100</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        (tmp_path / &quot;.gitignore&quot;).write_text(&quot;ignored/\n*.log\n&quot;)</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>101</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        (tmp_path / &quot;keep.py&quot;).write_text(&quot;print('hi')\n&quot;)</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>102</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        (tmp_path / &quot;debug.log&quot;).write_text(&quot;noise\n&quot;)</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>103</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        (tmp_path / &quot;ignored&quot;).mkdir()</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>104</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        (tmp_path / &quot;ignored&quot; / &quot;junk.py&quot;).write_text(&quot;x = 1\n&quot;)</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>105</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;> </div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>106</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        results = scan_files(target=tmp_path)</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>107</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        names = {r[0].name for r in results}</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>108</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        assert &quot;keep.py&quot; in names</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>109</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        assert &quot;debug.log&quot; not in names</div></div>
      <div class=&quot;row add&quot;><div class=&quot;gutter&quot;>110</div><div class=&quot;sign&quot;>+</div><div class=&quot;code&quot;>        assert &quot;junk.py&quot; not in names</div></div>
    </div>
    <div class=&quot;notes&quot;>
      <div class=&quot;ann pass&quot;>
        <div class=&quot;anchor&quot;>test as a whole</div>
        <div class=&quot;head&quot;>Pass &amp;mdash; exercises the fixed path</div>
        <div class=&quot;body&quot;>
          <code>tmp_path</code> is a fresh directory with no <code>.git</code> anywhere upstream (pytest's tmp lives outside any project repo by default). It covers both a glob (<code>*.log</code>) and a directory (<code>ignored/</code>), and asserts both positive (<code>keep.py</code> kept) and negative (two paths excluded). This test would have failed against the pre-fix code.
        </div>
      </div>
      <div class=&quot;ann nit&quot;>
        <div class=&quot;anchor&quot;>coverage gap</div>
        <div class=&quot;head&quot;>Nit &amp;mdash; consider one more case</div>
        <div class=&quot;body&quot;>
          <code>load_gitignore_specs</code> supports nested <code>.gitignore</code> files (it walks the whole subtree and prefixes patterns with the relative directory). The new test only places a <code>.gitignore</code> at the root of <code>tmp_path</code>. A second test with <code>tmp_path / &quot;sub&quot; / &quot;.gitignore&quot;</code> and a file inside <code>sub/</code> would lock in the nested-spec behavior under the new fallback path &amp;mdash; that's the part most likely to regress in a future refactor.
        </div>
      </div>
      <div class=&quot;ann warn&quot;>
        <div class=&quot;anchor&quot;>latent assumption</div>
        <div class=&quot;head&quot;>Concern &amp;mdash; test would silently fail if pytest tmp ever lived under a <code>.git</code></div>
        <div class=&quot;body&quot;>
          The test depends on pytest's <code>tmp_path</code> not having a <code>.git</code> ancestor. That's true today on every CI runner I'm aware of, and almost always true locally, but it's an implicit assumption. Adding an explicit <code>assert find_git_root(start=tmp_path) is None</code> at the top of the test would make the assumption visible and produce a clearer failure if it ever broke.
        </div>
      </div>
    </div>
  </div>
</div>

<h2>Findings summary</h2>

<div class=&quot;findings&quot;>

  <div class=&quot;finding pass&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Pass</span>
        <span class=&quot;title&quot;>Fix is correct and minimally scoped</span>
      </div>
      <span class=&quot;where&quot;>scanner.py:107&amp;ndash;111, 131&amp;ndash;132</span>
      <p>The same root value flows through spec construction and per-file relativization. No new code paths added; behavior with a real git root is byte-for-byte unchanged.</p>
    </div>
  </div>

  <div class=&quot;finding pass&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Pass</span>
        <span class=&quot;title&quot;>Variable rename improves accuracy</span>
      </div>
      <span class=&quot;where&quot;>scanner.py:107</span>
      <p><code>git_root</code> &amp;rarr; <code>gitignore_root</code> reflects the broadened semantics. Reading the new code, the intent is clearer than the old code's intent was.</p>
    </div>
  </div>

  <div class=&quot;finding pass&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Pass</span>
        <span class=&quot;title&quot;>New test exercises the bug</span>
      </div>
      <span class=&quot;where&quot;>test_scanner.py:99&amp;ndash;110</span>
      <p>Test fails on pre-fix code, passes on post-fix code, covers both a glob pattern and a directory pattern.</p>
    </div>
  </div>

  <div class=&quot;finding warn&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Concern</span>
        <span class=&quot;title&quot;>Pre-existing: an unrelated upstream <code>.git</code> hijacks the root</span>
      </div>
      <span class=&quot;where&quot;>scanner.py:109 (and find_git_root)</span>
      <p><code>find_git_root</code> walks to <code>/</code>. With <code>~/.git</code> (dotfiles), running toks on <code>~/projects/whatever</code> picks up the dotfiles repo and triggers <code>os.walk(~)</code> in <code>load_gitignore_specs</code>. Not introduced by this PR. Worth a follow-up to either bound the upward walk (e.g. stop at the user's home or filesystem boundary) or treat that case the way you treat &amp;ldquo;no git root&amp;rdquo; here.</p>
    </div>
  </div>

  <div class=&quot;finding warn&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Concern</span>
        <span class=&quot;title&quot;>Pre-existing: gitignored directories are still walked into</span>
      </div>
      <span class=&quot;where&quot;>scanner.py:scan_files loop</span>
      <p>The PR description specifically calls out <code>.venv</code>. With this fix, <code>.venv</code>'s files are correctly excluded from results, but <code>os.walk</code> still descends through every <code>site-packages</code> file and stats it. On a fresh venv that's tens of thousands of stats. Pruning <code>dirnames</code> against <code>gitignore_spec</code> before recursing would address it. Out of scope for this PR; flag as follow-up.</p>
    </div>
  </div>

  <div class=&quot;finding nit&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Nit</span>
        <span class=&quot;title&quot;><code>load_gitignore_specs</code> parameter still named <code>git_root</code></span>
      </div>
      <span class=&quot;where&quot;>scanner.py:73</span>
      <p>You renamed the variable in the caller; the parameter in the callee is now slightly misleading (callers may pass a non-git path). Trivial follow-up rename.</p>
    </div>
  </div>

  <div class=&quot;finding nit&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Nit</span>
        <span class=&quot;title&quot;><code>and gitignore_root</code> guard is redundant</span>
      </div>
      <span class=&quot;where&quot;>scanner.py:131</span>
      <p>After the fix, the two flags are coupled. Harmless but no longer informative.</p>
    </div>
  </div>

  <div class=&quot;finding nit&quot;>
    <div class=&quot;stripe&quot;></div>
    <div class=&quot;body&quot;>
      <div class=&quot;head&quot;>
        <span class=&quot;badge&quot;>Nit</span>
        <span class=&quot;title&quot;>One more test would lock in nested-<code>.gitignore</code> behavior on the fallback path</span>
      </div>
      <span class=&quot;where&quot;>test_scanner.py</span>
      <p>Add a test where <code>tmp_path/sub/.gitignore</code> defines a pattern and a file inside <code>sub/</code> matches it &amp;mdash; verifies <code>load_gitignore_specs</code>'s prefix logic still works when the root is the target rather than a real git root.</p>
    </div>
  </div>

</div>

<h2>Behavioral coverage matrix</h2>
<p>Cases the implementation can encounter and where each is covered:</p>

<table class=&quot;cov&quot;>
  <thead>
    <tr>
      <th>Scenario</th>
      <th>Pre-PR behavior</th>
      <th>Post-PR behavior</th>
      <th>Test coverage</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Target is inside a git repo, has <code>.gitignore</code></td>
      <td>Honored</td>
      <td>Honored (unchanged)</td>
      <td><span class=&quot;yes&quot;>Yes</span> &amp;mdash; existing fixture-based tests</td>
    </tr>
    <tr>
      <td>Target has <code>.gitignore</code>, no <code>.git</code> upstream</td>
      <td><span class=&quot;no&quot;>Silently ignored</span> (the bug)</td>
      <td>Honored via target-as-root fallback</td>
      <td><span class=&quot;yes&quot;>Yes</span> &amp;mdash; new test</td>
    </tr>
    <tr>
      <td>Target has <code>.gitignore</code> at root <em>and</em> a nested <code>sub/.gitignore</code>, no <code>.git</code> upstream</td>
      <td>Both ignored (bug)</td>
      <td>Both should be honored (relies on <code>load_gitignore_specs</code> walking from <code>target</code>)</td>
      <td><span class=&quot;no&quot;>Not directly tested</span></td>
    </tr>
    <tr>
      <td>Target with no <code>.gitignore</code> and no <code>.git</code></td>
      <td>No filtering</td>
      <td>No filtering (<code>load_gitignore_specs</code> returns <code>None</code>)</td>
      <td>Implicit via <code>test_empty_dir</code></td>
    </tr>
    <tr>
      <td>Target with <code>no_gitignore=True</code></td>
      <td>Skipped</td>
      <td>Skipped (unchanged)</td>
      <td>Not directly tested (pre-existing gap)</td>
    </tr>
    <tr>
      <td>Target inside an unrelated upstream <code>.git</code> (e.g. <code>~/.git</code>)</td>
      <td><span class=&quot;no&quot;>Walks entire ancestor tree</span></td>
      <td><span class=&quot;no&quot;>Same &amp;mdash; not addressed</span></td>
      <td>Not tested</td>
    </tr>
  </tbody>
</table>

<h2>Recommendation</h2>
<p>Ship this PR as-is. The fix is correct, well-named, and has a test that would catch a regression. The two pre-existing concerns (upstream-<code>.git</code> hijacking and per-file rather than per-directory filtering) are worth filing as follow-ups but should not block this merge &amp;mdash; widening scope would dilute a clean, easy-to-review change.</p>

</body>
</html>
" height="700" width="100%" style="border:1px solid #d0d7de;border-radius:6px;margin:1.5em 0"></iframe>

*The HTML PR review. Verdict bar, severity-tagged findings, annotated diff. Polished. But the markdown version below has the same substance.*

<iframe srcdoc="<!doctype html>
<html lang=&quot;en&quot;>
<head>
<meta charset=&quot;utf-8&quot;>
<title>Rendered markdown</title>
<style>
  body { font: 15px/1.6 -apple-system, BlinkMacSystemFont, &quot;Segoe UI&quot;, Roboto, sans-serif;
         color: #1f2328; background: #ffffff; max-width: 760px; margin: 0 auto; padding: 24px; }
  h1 { font-size: 24px; margin-top: 0; border-bottom: 1px solid #d0d7de; padding-bottom: 8px; }
  h2 { font-size: 20px; margin-top: 28px; border-bottom: 1px solid #d0d7de; padding-bottom: 6px; }
  h3 { font-size: 16px; margin-top: 22px; }
  code { background: #f6f8fa; padding: 1px 5px; border-radius: 4px; font-size: 13px;
         font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace; }
  pre { background: #f6f8fa; padding: 12px 14px; border-radius: 6px; overflow-x: auto;
        font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace; font-size: 13px;
        line-height: 1.45; }
  pre code { background: transparent; padding: 0; font-size: inherit; }
  blockquote { border-left: 3px solid #d0d7de; padding-left: 12px; color: #59636e; margin: 12px 0; }
  table { border-collapse: collapse; margin: 12px 0; font-size: 14px; }
  th, td { border: 1px solid #d0d7de; padding: 6px 12px; text-align: left; vertical-align: top; }
  th { background: #f6f8fa; }
  hr { border: none; border-top: 1px solid #d0d7de; margin: 28px 0; }
  a { color: #0969da; }
</style>
</head>
<body>
<h1>PR Review — <code>toks</code> c6d70f9</h1>
<p><strong>Title:</strong> Respect .gitignore when target has no .git directory<br />
<strong>Author:</strong> Corey Gallon<br />
<strong>Date:</strong> 2026-04-29<br />
<strong>Files touched:</strong> <code>src/toks/scanner.py</code> (+4 / -4), <code>tests/test_scanner.py</code> (+13 / -0)</p>
<hr />
<h2>TL;DR</h2>
<p><strong>Verdict: Approve with two non-blocking notes.</strong></p>
<p>The fix is minimal, correct for the bug it targets, and well-tested. The rename from <code>git_root</code> to <code>gitignore_root</code> is the right framing — the variable is now &quot;the root we're matching gitignore patterns relative to,&quot; and that name finally tells the truth.</p>
<p>The two notes worth flagging before merge:</p>
<ol>
<li>The fallback intentionally narrows discovery: <code>.gitignore</code> files <strong>above</strong> the target are no longer consulted when no <code>.git</code> is present (they never were, but the fallback formalizes that). Worth a one-line docstring note so future-you doesn't re-litigate it.</li>
<li>The new test is mildly fragile to ancestor-<code>.git</code> pollution from <code>tmp_path</code>. In practice fine; a single <code>no_gitignore=False</code> is implicit and may want a sibling assertion to lock the contract.</li>
</ol>
<p>Neither blocks the merge.</p>
<hr />
<h2>Severity legend</h2>
<table>
<thead>
<tr>
<th>Tag</th>
<th>Meaning</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>[CRITICAL]</code></td>
<td>Blocks merge. Correctness, security, data loss.</td>
</tr>
<tr>
<td><code>[HIGH]</code></td>
<td>Should fix before merge. Likely to bite users.</td>
</tr>
<tr>
<td><code>[MEDIUM]</code></td>
<td>Address before merge if cheap, otherwise track.</td>
</tr>
<tr>
<td><code>[LOW]</code></td>
<td>Nit / polish. Optional.</td>
</tr>
<tr>
<td><code>[POSITIVE]</code></td>
<td>Worth calling out — got this right.</td>
</tr>
</tbody>
</table>
<p>This review found: <strong>0 CRITICAL · 0 HIGH · 2 MEDIUM · 3 LOW · 2 POSITIVE.</strong></p>
<hr />
<h2>Bracing the gitignore / path-resolution logic</h2>
<p>I want to walk this carefully because the bug is exactly the kind of thing that hides in path semantics. Here's the model after the change:</p>
<pre><code>target = the directory the user pointed toks at
git_root = nearest ancestor containing .git (or None)
gitignore_root = git_root if git_root else target   # NEW
</code></pre>
<p><code>gitignore_root</code> is then used for two things:</p>
<ol>
<li><strong>As the walk root for <code>.gitignore</code> discovery</strong> — <code>load_gitignore_specs(git_root=...)</code> does <code>os.walk(git_root)</code> and concatenates patterns from every <code>.gitignore</code> it finds, prefixed by their relative path under that root.</li>
<li><strong>As the basis for relativization at match time</strong> — <code>file_path.relative_to(gitignore_root)</code> produces the path string passed to <code>pathspec.match_file</code>.</li>
</ol>
<p>These two uses must use the <strong>same</strong> root for matching to be correct. The PR keeps them in lockstep — that's the load-bearing invariant, and it's preserved.</p>
<h3>Walking the four cases</h3>
<table>
<thead>
<tr>
<th>Case</th>
<th>git found?</th>
<th><code>git_root</code></th>
<th><code>gitignore_root</code> (after)</th>
<th>Discovery walks</th>
<th>Relativizes against</th>
<th>Correct?</th>
</tr>
</thead>
<tbody>
<tr>
<td>Target inside a git repo</td>
<td>yes</td>
<td><code>/repo</code></td>
<td><code>/repo</code></td>
<td><code>/repo</code></td>
<td><code>/repo</code></td>
<td>yes (unchanged)</td>
</tr>
<tr>
<td>Target IS the git root</td>
<td>yes</td>
<td><code>target</code></td>
<td><code>target</code></td>
<td><code>target</code></td>
<td><code>target</code></td>
<td>yes (unchanged)</td>
</tr>
<tr>
<td>Target has its own <code>.gitignore</code>, no <code>.git</code> anywhere</td>
<td>no</td>
<td><code>None</code></td>
<td><code>target</code></td>
<td><code>target</code></td>
<td><code>target</code></td>
<td><strong>yes — newly fixed</strong></td>
</tr>
<tr>
<td>Target has no <code>.gitignore</code>, no <code>.git</code></td>
<td>no</td>
<td><code>None</code></td>
<td><code>target</code></td>
<td><code>target</code> (yields no patterns)</td>
<td><code>target</code></td>
<td>yes — <code>load_gitignore_specs</code> returns <code>None</code>, so the guard <code>if gitignore_spec and gitignore_root:</code> short-circuits</td>
</tr>
</tbody>
</table>
<p>The fourth row is the one I'd want a reviewer to verify by reading. <code>load_gitignore_specs</code> returns <code>None</code> when <code>patterns</code> is empty, the assignment becomes <code>gitignore_spec = None</code>, and the per-file guard skips matching. No regression.</p>
<h3>What the fallback does NOT do</h3>
<p>There's one behavior it would be easy to assume but isn't true: <strong>the fallback does not search upward for <code>.gitignore</code> files when there's no <code>.git</code>.</strong> If the layout is</p>
<pre><code>/parent/
  .gitignore        # contains &amp;quot;*.log&amp;quot;
  child/
    debug.log
</code></pre>
<p>…and the user runs <code>toks /parent/child</code>, <code>debug.log</code> will be scanned. <code>find_git_root</code> returns <code>None</code> (no ancestor has <code>.git</code>), <code>gitignore_root</code> falls back to <code>child</code>, and <code>os.walk(child)</code> never sees <code>/parent/.gitignore</code>.</p>
<p>This is consistent with the commit message (&quot;fall back to the target directory itself&quot;) and consistent with how git itself behaves (no <code>.git</code> → no project boundary → no upward <code>.gitignore</code> chain). I'd just like one line in the docstring saying so, because the next person to think about this will think about it again otherwise.</p>
<hr />
<h2>The diff, annotated</h2>
<p>Below: each hunk, followed by per-line notes. Annotations cite line numbers from the <em>new</em> file.</p>
<h3>Hunk 1 — <code>scanner.py</code> lines 104-112</h3>
<pre><code class=&quot;language-python&quot;>   target = target.resolve()
   if not target.is_dir():
       raise ValueError(f&amp;quot;Not a directory: {target}&amp;quot;)

   gitignore_spec = None
-  git_root = None
+  gitignore_root = None                                                       # ← (A)
   if not no_gitignore:
       git_root = find_git_root(start=target)
-      if git_root:
-          gitignore_spec = load_gitignore_specs(git_root=git_root, target=target)
+      gitignore_root = git_root if git_root else target                       # ← (B)
+      gitignore_spec = load_gitignore_specs(git_root=gitignore_root, target=target)  # ← (C)
</code></pre>
<table>
<thead>
<tr>
<th>Mark</th>
<th>Annotation</th>
<th>Severity</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>A</strong></td>
<td>Rename is the right call — this variable now means &quot;the root we relativize against,&quot; not &quot;the git repo root.&quot; Name follows semantics.</td>
<td><code>[POSITIVE]</code></td>
</tr>
<tr>
<td><strong>B</strong></td>
<td>Ternary is readable. Could equivalently be <code>gitignore_root = git_root or target</code>, which is shorter and idiomatic Python. Style preference; either is fine.</td>
<td><code>[LOW]</code></td>
</tr>
<tr>
<td><strong>C</strong></td>
<td>The <code>target=target</code> keyword arg is now unused inside <code>load_gitignore_specs</code> — <code>target_resolved = target.resolve()</code> is computed and never read. Pre-existing dead code, but this PR is the natural moment to either delete the parameter or use it. See finding <strong>F-3</strong> below.</td>
<td><code>[MEDIUM]</code></td>
</tr>
</tbody>
</table>
<h3>Hunk 2 — <code>scanner.py</code> lines 128-135</h3>
<pre><code class=&quot;language-python&quot;>           if file_path.is_symlink() and file_path.is_dir():
               continue

-          if gitignore_spec and git_root:
-              rel = file_path.relative_to(git_root)
+          if gitignore_spec and gitignore_root:                               # ← (D)
+              rel = file_path.relative_to(gitignore_root)                     # ← (E)
               if gitignore_spec.match_file(str(rel)):
                   continue
</code></pre>
<table>
<thead>
<tr>
<th>Mark</th>
<th>Annotation</th>
<th>Severity</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>D</strong></td>
<td>Guard updated in lockstep with the rename. The two checks (<code>gitignore_spec</code>, <code>gitignore_root</code>) are now both truthy iff we successfully built a spec — <code>gitignore_spec</code> alone is sufficient (since spec is only built when root is set), but the redundant guard is defensive and harmless.</td>
<td><code>[POSITIVE]</code></td>
</tr>
<tr>
<td><strong>E</strong></td>
<td><code>relative_to(gitignore_root)</code> is safe: <code>file_path</code> comes from <code>os.walk(target)</code>, and <code>gitignore_root</code> is either <code>target</code> or an ancestor of it, so <code>file_path</code> is always under <code>gitignore_root</code>. No <code>ValueError</code> risk.</td>
<td><code>[POSITIVE]</code></td>
</tr>
</tbody>
</table>
<h3>Hunk 3 — <code>tests/test_scanner.py</code> lines 99-110</h3>
<pre><code class=&quot;language-python&quot;>+  def test_gitignore_respected_without_git_dir(self, tmp_path):
+      (tmp_path / &amp;quot;.gitignore&amp;quot;).write_text(&amp;quot;ignored/\n*.log\n&amp;quot;)
+      (tmp_path / &amp;quot;keep.py&amp;quot;).write_text(&amp;quot;print('hi')\n&amp;quot;)
+      (tmp_path / &amp;quot;debug.log&amp;quot;).write_text(&amp;quot;noise\n&amp;quot;)
+      (tmp_path / &amp;quot;ignored&amp;quot;).mkdir()
+      (tmp_path / &amp;quot;ignored&amp;quot; / &amp;quot;junk.py&amp;quot;).write_text(&amp;quot;x = 1\n&amp;quot;)
+
+      results = scan_files(target=tmp_path)                                   # ← (F)
+      names = {r[0].name for r in results}
+      assert &amp;quot;keep.py&amp;quot; in names
+      assert &amp;quot;debug.log&amp;quot; not in names                                         # ← (G)
+      assert &amp;quot;junk.py&amp;quot; not in names                                           # ← (H)
</code></pre>
<table>
<thead>
<tr>
<th>Mark</th>
<th>Annotation</th>
<th>Severity</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>F</strong></td>
<td>Test exercises the regression directly: <code>tmp_path</code> is system tmp on Linux/macOS and typically has no ancestor <code>.git</code>. Fragility note in <strong>F-1</strong> below.</td>
<td>—</td>
</tr>
<tr>
<td><strong>G</strong></td>
<td>Covers the file-glob case (<code>*.log</code>).</td>
<td><code>[POSITIVE]</code></td>
</tr>
<tr>
<td><strong>H</strong></td>
<td>Covers the directory case (<code>ignored/</code>). Both major gitignore pattern shapes get coverage.</td>
<td><code>[POSITIVE]</code></td>
</tr>
</tbody>
</table>
<p>The test asserts the <em>positive</em> (<code>keep.py</code> in) and the <em>negatives</em> (<code>debug.log</code>, <code>junk.py</code> not in). That's the right shape — a test that only asserted exclusions could pass with <code>scan_files</code> returning <code>[]</code>.</p>
<hr />
<h2>Findings</h2>
<h3><code>[MEDIUM]</code> F-1 — Test is silently dependent on <code>tmp_path</code> having no ancestor <code>.git</code></h3>
<p><code>find_git_root</code> walks upward from <code>target.resolve()</code> until it hits the filesystem root, looking for any <code>.git</code>. If a developer's test environment puts <code>tmp_path</code> somewhere under a git checkout (rare but possible — custom <code>tmpdir</code> configs, certain CI sandboxes, network-mounted home dirs with stray <code>.git</code> symlinks), the test would skip the new code path entirely. It would still likely pass because of how the patterns happen to be structured, but it would no longer be testing what its name claims.</p>
<p><strong>Suggested fix:</strong> Either (a) explicitly verify no ancestor has <code>.git</code> at the start of the test, or (b) make the intent unambiguous by also asserting via a second call with <code>no_gitignore=True</code> that the included set differs:</p>
<pre><code class=&quot;language-python&quot;>results_no_ignore = scan_files(target=tmp_path, no_gitignore=True)
no_ignore_names = {r[0].name for r in results_no_ignore}
assert &amp;quot;debug.log&amp;quot; in no_ignore_names  # confirms gitignore did the filtering
</code></pre>
<p>That second assertion locks the contract: &quot;filtering happened <em>because of</em> gitignore handling,&quot; not &quot;filtering happened, somehow.&quot;</p>
<h3><code>[MEDIUM]</code> F-2 — Behavior of fallback should be documented</h3>
<p>The docstring of <code>scan_files</code> doesn't mention gitignore handling at all today. With this change, the rule &quot;we'll honor a <code>.gitignore</code> in target even without <code>.git</code>&quot; is now part of the contract. One line is enough:</p>
<pre><code class=&quot;language-python&quot;>&amp;quot;&amp;quot;&amp;quot;Scan a directory for files, returning (path, mime_type, file_size) tuples.

…

When no_gitignore is False (default), .gitignore files are honored. The
gitignore root is the nearest ancestor containing .git, or the target
itself if no such ancestor exists. .gitignore files above the gitignore
root are not consulted.
&amp;quot;&amp;quot;&amp;quot;
</code></pre>
<h3><code>[LOW]</code> F-3 — Unused parameter in <code>load_gitignore_specs</code></h3>
<pre><code class=&quot;language-python&quot;>def load_gitignore_specs(*, git_root: Path, target: Path) -&amp;gt; pathspec.PathSpec | None:
    patterns: list[str] = []
    target_resolved = target.resolve()  # ← never read
    ...
</code></pre>
<p><code>target</code> is no longer used inside this function. Pre-existing, not introduced by this PR — but the PR is the natural pass-by for it. Either remove the parameter and the dead line, or use it (e.g., to skip <code>.gitignore</code> files outside the target subtree if you wanted to make scoping stricter — though I'd argue you don't, because git itself doesn't).</p>
<p>If removing: also rename <code>git_root</code> → <code>root</code> while you're there, since the parameter is no longer git-specific in concept.</p>
<h3><code>[LOW]</code> F-4 — Style: <code>git_root or target</code> over the ternary</h3>
<pre><code class=&quot;language-python&quot;>gitignore_root = git_root if git_root else target
# vs
gitignore_root = git_root or target
</code></pre>
<p><code>Path</code> instances are always truthy, and <code>find_git_root</code> returns <code>None</code> or a <code>Path</code>, so the short form is both safe and idiomatic. Pure preference.</p>
<h3><code>[LOW]</code> F-5 — <code>find_git_root</code> accepts <code>.git</code> as either file or directory; commit message says &quot;directory&quot;</h3>
<p><code>find_git_root</code> checks <code>(current / &quot;.git&quot;).exists()</code>, which is true for both directories and files. Worktrees use a <code>.git</code> <em>file</em> (containing <code>gitdir: ...</code>). The commit message says &quot;no <code>.git</code> directory,&quot; but the code does the right thing for worktrees too. This is a wording nit on the commit message, not a code issue. The change correctly does not regress worktree handling.</p>
<h3><code>[POSITIVE]</code> F-6 — Minimal diff</h3>
<p>Four-line change in production code, two of them pure renames, plus a focused test. Doesn't touch unrelated logic. Doesn't introduce new abstractions. The kind of fix that ages well.</p>
<h3><code>[POSITIVE]</code> F-7 — Test assertions cover both pattern shapes</h3>
<p><code>*.log</code> (file glob) and <code>ignored/</code> (directory) are the two pattern syntaxes most users care about, and both are exercised. A <code>.gitignore</code> parser regression in either would fail this test.</p>
<hr />
<h2>Suggested commit-message tweak (optional)</h2>
<blockquote>
<p>Respect .gitignore when target has no .git directory <strong>or worktree marker</strong></p>
</blockquote>
<p>Tiny edit; keeps the message accurate for the worktree-file case the code already handles.</p>
<hr />
<h2>Pre-merge checklist</h2>
<ul>
<li>[ ] Add the docstring note from <strong>F-2</strong> (10 seconds).</li>
<li>[ ] Optional: tighten the test per <strong>F-1</strong> (one extra <code>no_gitignore=True</code> call).</li>
<li>[ ] Optional: address <strong>F-3</strong> in a follow-up cleanup commit.</li>
<li>[ ] No security implications. No performance regression: when <code>git_root</code> is <code>None</code> and target has no <code>.gitignore</code>, <code>load_gitignore_specs</code> walks the target tree once and returns <code>None</code>. That walk is bounded by the same tree the main <code>os.walk</code> traverses anyway, so worst case is one extra traversal of a directory the user already chose to scan.</li>
</ul>
<p>Ready to merge after F-2.</p></body></html>" height="700" width="100%" style="border:1px solid #d0d7de;border-radius:6px;margin:1.5em 0"></iframe>

*The same review, in markdown. Same content. Roughly 30% cheaper. The cost of "looks nicer" is a real number, and that's it.*

I also asked Claude to make a semantic edit to each artifact (add a "Test coverage assessment" section) and looked at the resulting diffs. HTML's was 5,658 bytes / 93 lines; markdown's was 4,618 bytes / 55 lines. HTML's structural overhead (1.2x more bytes than markdown for the same logical edit) is real but modest, because both formats already had a fair bit of structure to begin with.

So: markdown does the job. HTML is prettier and costs about 50% more. Worth it for a review template you'll reuse across the team and reference repeatedly; probably not for a one-off you'll glance at and close.

## Rate limiter explainer: where the markdown actually won on substance

> **Methodology note:** same as above, verbatim ablation only. The 1.44x cost ratio is from a HTML-affording prompt and would likely be smaller under a neutral one.

**Thariq's prompt (lightly adapted):** *"I don't understand how slowapi's rate limiter actually works. Read the relevant code and produce a single HTML explainer page: a diagram of the rate-limiting flow, the 3-4 key code snippets annotated, and a 'gotchas' section..."*

**Substrate:** `slowapi` core (`extension.py`, `middleware.py`, `wrappers.py`), about 12K tokens, ~1,200 lines of Python.

**Cost data:**

| Format | Artifact tokens | Output tokens | Generation time | Cost |
|---|---|---|---|---|
| HTML | 8,700 | 14,713 | 191 s | $0.79 |
| MD   | 3,629 | 7,915  | 118 s | $0.55 |

Ratios: 2.40x artifact tokens, 1.86x output tokens, 1.62x time, 1.44x cost.

Both artifacts are real explainers. The bit I wasn't expecting: **both used Mermaid for the flow diagram.** The markdown one wraps it in a `` ```mermaid `` fence; the HTML one loads the Mermaid CDN and embeds the same chart definition. GitHub renders the markdown version as a real diagram. So does VS Code. The "HTML can show diagrams and markdown can't" intuition that does some quiet work in Thariq's argument is mostly gone for technical writing in 2026.

<iframe srcdoc="<!DOCTYPE html>
<html lang=&quot;en&quot;>
<head>
<meta charset=&quot;utf-8&quot;>
<title>How slowapi's rate limiter actually works</title>
<script src=&quot;https://cdn.jsdelivr.net/npm/mermaid@10/dist/mermaid.min.js&quot;></script>
<script>
  mermaid.initialize({ startOnLoad: true, theme: 'neutral', flowchart: { useMaxWidth: true, htmlLabels: true } });
</script>
<style>
  :root {
    --fg: #1a1a1a;
    --muted: #555;
    --bg: #fafafa;
    --panel: #ffffff;
    --border: #d8d8d8;
    --accent: #0b5fff;
    --code-bg: #0f1115;
    --code-fg: #e6e6e6;
    --kw: #ff7b72;
    --str: #a5d6ff;
    --com: #8b949e;
    --num: #d2a8ff;
    --hilite: #fff7c2;
    --warn: #b54708;
    --warn-bg: #fff7ed;
    --warn-border: #fdba74;
  }
  html { box-sizing: border-box; }
  *, *:before, *:after { box-sizing: inherit; }
  body {
    margin: 0;
    font-family: -apple-system, BlinkMacSystemFont, &quot;Segoe UI&quot;, Roboto, Helvetica, Arial, sans-serif;
    color: var(--fg);
    background: var(--bg);
    line-height: 1.55;
    font-size: 16px;
  }
  .wrap {
    max-width: 980px;
    margin: 0 auto;
    padding: 36px 28px 80px;
  }
  h1 {
    font-size: 28px;
    margin: 0 0 4px;
    letter-spacing: -0.01em;
  }
  .subtitle {
    color: var(--muted);
    margin: 0 0 28px;
    font-size: 15px;
  }
  h2 {
    font-size: 21px;
    margin: 36px 0 12px;
    padding-bottom: 6px;
    border-bottom: 1px solid var(--border);
    letter-spacing: -0.01em;
  }
  h3 {
    font-size: 16px;
    margin: 24px 0 8px;
    color: #222;
  }
  p { margin: 0 0 12px; }
  .lede {
    background: var(--panel);
    border-left: 3px solid var(--accent);
    padding: 14px 18px;
    border-radius: 4px;
    margin-bottom: 20px;
  }
  .panel {
    background: var(--panel);
    border: 1px solid var(--border);
    border-radius: 6px;
    padding: 16px 18px;
    margin: 12px 0 20px;
  }
  .diagram {
    background: var(--panel);
    border: 1px solid var(--border);
    border-radius: 6px;
    padding: 14px;
    overflow-x: auto;
  }
  pre {
    background: var(--code-bg);
    color: var(--code-fg);
    padding: 14px 16px;
    border-radius: 6px;
    overflow-x: auto;
    font-size: 13.5px;
    font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
    line-height: 1.5;
    margin: 8px 0 0;
  }
  pre .kw { color: var(--kw); }
  pre .str { color: var(--str); }
  pre .com { color: var(--com); font-style: italic; }
  pre .num { color: var(--num); }
  pre .hi {
    background: rgba(255, 247, 194, 0.18);
    display: inline-block;
    width: 100%;
  }
  code.inline {
    background: #eef0f3;
    padding: 1px 6px;
    border-radius: 3px;
    font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
    font-size: 0.9em;
  }
  .annot {
    margin: 0 0 8px;
    color: var(--muted);
    font-size: 14.5px;
  }
  .snippet {
    margin-bottom: 22px;
  }
  .snippet-title {
    font-weight: 600;
    font-size: 15px;
    margin-bottom: 4px;
  }
  .file-tag {
    font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace;
    font-size: 12px;
    color: var(--muted);
    margin-left: 6px;
    font-weight: normal;
  }
  ol.gotchas, ul.gotchas { padding-left: 20px; }
  ol.gotchas li, ul.gotchas li { margin-bottom: 14px; }
  .gotcha {
    border-left: 3px solid var(--warn-border);
    background: var(--warn-bg);
    padding: 10px 14px;
    border-radius: 4px;
    margin: 10px 0;
  }
  .gotcha .label {
    color: var(--warn);
    font-weight: 600;
    font-size: 12.5px;
    letter-spacing: 0.04em;
    text-transform: uppercase;
    margin-bottom: 4px;
  }
  .pill {
    display: inline-block;
    padding: 1px 8px;
    border-radius: 999px;
    background: #eef2ff;
    color: #1e3a8a;
    font-size: 12px;
    margin-right: 6px;
    font-weight: 600;
  }
  table {
    border-collapse: collapse;
    width: 100%;
    margin: 8px 0 14px;
    font-size: 14.5px;
  }
  th, td {
    border: 1px solid var(--border);
    padding: 8px 10px;
    text-align: left;
    vertical-align: top;
  }
  th { background: #f3f4f6; font-weight: 600; }
</style>
</head>
<body>
<div class=&quot;wrap&quot;>

<h1>How slowapi's rate limiter actually works</h1>
<p class=&quot;subtitle&quot;>A one-shot tour of <code class=&quot;inline&quot;>extension.py</code>, <code class=&quot;inline&quot;>middleware.py</code>, and <code class=&quot;inline&quot;>wrappers.py</code>.</p>

<div class=&quot;lede&quot;>
  <p><strong>The shape in one paragraph.</strong> slowapi is a thin orchestration layer over the <code class=&quot;inline&quot;>limits</code> library. It collects rate-limit declarations from three places (decorators, app-level <em>application</em> limits, app-level <em>default</em> limits), and on each request it builds a list of which of those apply, calls <code class=&quot;inline&quot;>limits.hit()</code> on each in turn against a configured backend (memory / Redis / etc.), and either lets the request through or raises <code class=&quot;inline&quot;>RateLimitExceeded</code>. The check can be triggered from a decorator wrapper or from a Starlette middleware -- they are two entry points into the same core function, <code class=&quot;inline&quot;>_check_request_limit</code>.</p>
</div>

<h2>Rate-limiting flow</h2>

<div class=&quot;diagram&quot;>
<div class=&quot;mermaid&quot;>
flowchart TD
  Req([Incoming HTTP request])
  Req --> Entry{&quot;Entry point&quot;}

  Entry -->|&quot;Middleware path<br/>(SlowAPIMiddleware /<br/>SlowAPIASGIMiddleware)&quot;| MW[Find route handler<br/>via app.routes]
  Entry -->|&quot;Decorator path<br/>(@limiter.limit / @limiter.shared_limit)&quot;| Dec[Decorator wrapper<br/>extracts request from args/kwargs]

  MW --> Exempt{&quot;_should_exempt?<br/>handler missing OR<br/>name in _exempt_routes OR<br/>name in _route_limits&quot;}
  Exempt -->|yes| Pass[Pass through, no check]
  Exempt -->|no| CallCore[&quot;_check_request_limit<br/>(in_middleware=True)&quot;]

  Dec --> Flag{&quot;request.state._rate_limiting_complete<br/>already True?&quot;}
  Flag -->|yes| RunHandler1[Run handler]
  Flag -->|no| CallCoreD[&quot;_check_request_limit<br/>(in_middleware=False)&quot;]
  CallCoreD --> SetFlag[Set _rate_limiting_complete = True]
  SetFlag --> RunHandler1

  CallCore --> Build
  CallCoreD --> Build

  Build[&quot;Build all_limits list:<br/>• application_limits (only if in_middleware)<br/>• route_limits + dynamic_route_limits (only if NOT in_middleware)<br/>• default_limits (unless route has override_defaults=True)&quot;]
  Build --> StorageDead{&quot;_storage_dead<br/>AND fallback_limiter?&quot;}
  StorageDead -->|yes| Fallback[Use _in_memory_fallback limits<br/>via _fallback_limiter]
  StorageDead -->|no| Eval

  Fallback --> Eval

  Eval[&quot;__evaluate_limits:<br/>for each Limit in all_limits&quot;]
  Eval --> ForEach{&quot;per-limit checks&quot;}
  ForEach -->|&quot;is_exempt(request)<br/>or method mismatch&quot;| Skip[Skip this limit]
  ForEach -->|&quot;otherwise&quot;| Hit[&quot;self.limiter.hit(<br/>limit, key_func(request), scope, cost=...)&quot;]
  Skip --> Eval

  Hit --> Allowed{&quot;hit returns True?&quot;}
  Allowed -->|yes, track smallest as<br/>limit_for_header| Eval
  Allowed -->|no| Fail[&quot;Set request.state.view_rate_limit<br/>raise RateLimitExceeded&quot;]

  Eval -->|&quot;all passed&quot;| Done[&quot;request.state.view_rate_limit<br/>= smallest limit seen&quot;]
  Done --> RunHandler2[Run handler / call_next]

  RunHandler1 --> Inject[_inject_headers / _inject_asgi_headers:<br/>X-RateLimit-Limit / Remaining / Reset / Retry-After]
  RunHandler2 --> Inject
  Fail --> ExcHandler[&quot;_rate_limit_exceeded_handler<br/>builds 429 JSON response&quot;]
  ExcHandler --> Inject
  Inject --> Resp([Response to client])
  Pass --> Resp
</div>
</div>

<h2>The four code paths that matter</h2>

<div class=&quot;snippet&quot;>
<div class=&quot;snippet-title&quot;>1. The core check: <code class=&quot;inline&quot;>_check_request_limit</code> <span class=&quot;file-tag&quot;>extension.py</span></div>
<p class=&quot;annot&quot;>Both entry points funnel here. The job of this function is to <em>assemble</em> the list of limits that apply to this request, then hand them to <code class=&quot;inline&quot;>__evaluate_limits</code>. The <code class=&quot;inline&quot;>in_middleware</code> flag is the pivot: it controls which buckets of limits are pulled in, so middleware and decorator don't double-count or miss each other.</p>
<pre><code><span class=&quot;kw&quot;>def</span> _check_request_limit(self, request, endpoint_func, in_middleware=<span class=&quot;num&quot;>True</span>):
    endpoint_url     = request[<span class=&quot;str&quot;>&quot;path&quot;</span>] <span class=&quot;kw&quot;>or</span> <span class=&quot;str&quot;>&quot;&quot;</span>
    endpoint_name    = <span class=&quot;kw&quot;>f</span><span class=&quot;str&quot;>&quot;{endpoint_func.__module__}.{endpoint_func.__name__}&quot;</span> <span class=&quot;kw&quot;>if</span> endpoint_func <span class=&quot;kw&quot;>else</span> <span class=&quot;str&quot;>&quot;&quot;</span>
    _endpoint_key    = endpoint_url <span class=&quot;kw&quot;>if</span> self._key_style == <span class=&quot;str&quot;>&quot;url&quot;</span> <span class=&quot;kw&quot;>else</span> endpoint_name

    <span class=&quot;com&quot;># Bail-outs: disabled, exempt, or a request_filter says skip.</span>
    <span class=&quot;kw&quot;>if</span> (<span class=&quot;kw&quot;>not</span> _endpoint_key <span class=&quot;kw&quot;>or not</span> self.enabled
        <span class=&quot;kw&quot;>or</span> endpoint_name <span class=&quot;kw&quot;>in</span> self._exempt_routes
        <span class=&quot;kw&quot;>or</span> any(fn() <span class=&quot;kw&quot;>for</span> fn <span class=&quot;kw&quot;>in</span> self._request_filters)):
        <span class=&quot;kw&quot;>return</span>

    limits, dynamic_limits = [], []
    <span class=&quot;kw&quot;>if not</span> in_middleware:                              <span class=&quot;com&quot;># decorator path only</span>
        limits         = self._route_limits.get(endpoint_name, [])
        dynamic_limits = [l <span class=&quot;kw&quot;>for</span> lg <span class=&quot;kw&quot;>in</span> self._dynamic_route_limits.get(endpoint_name, [])
                            <span class=&quot;kw&quot;>for</span> l <span class=&quot;kw&quot;>in</span> lg.with_request(request)]

    route_limits      = limits + dynamic_limits
    all_limits        = list(itertools.chain(*self._application_limits)) <span class=&quot;kw&quot;>if</span> in_middleware <span class=&quot;kw&quot;>else</span> []
    all_limits       += route_limits

    <span class=&quot;com&quot;># Defaults apply unless THIS route's limits all set override_defaults=True.</span>
    combined_defaults = all(<span class=&quot;kw&quot;>not</span> l.override_defaults <span class=&quot;kw&quot;>for</span> l <span class=&quot;kw&quot;>in</span> route_limits)
    <span class=&quot;kw&quot;>if</span> (<span class=&quot;kw&quot;>not</span> route_limits <span class=&quot;kw&quot;>or</span> combined_defaults):
        all_limits   += list(itertools.chain(*self._default_limits))

    self.__evaluate_limits(request, _endpoint_key, all_limits)</code></pre>
</div>

<div class=&quot;snippet&quot;>
<div class=&quot;snippet-title&quot;>2. The actual hit-or-miss: <code class=&quot;inline&quot;>__evaluate_limits</code> <span class=&quot;file-tag&quot;>extension.py</span></div>
<p class=&quot;annot&quot;>This is the hot loop. For each <code class=&quot;inline&quot;>Limit</code> it computes the bucket key (<code class=&quot;inline&quot;>key_func(request)</code> + scope), tracks the <em>smallest</em> limit seen so far for header reporting, and calls <code class=&quot;inline&quot;>self.limiter.hit(...)</code> -- this is the <code class=&quot;inline&quot;>limits</code> library's <code class=&quot;inline&quot;>RateLimiter</code>, the thing that actually increments the counter in Redis/memory. <code class=&quot;inline&quot;>hit</code> returning <code class=&quot;inline&quot;>False</code> means &quot;you're over&quot;; we capture the failed limit and break out so subsequent limits aren't decremented.</p>
<pre><code><span class=&quot;kw&quot;>def</span> __evaluate_limits(self, request, endpoint, limits):
    failed_limit = <span class=&quot;num&quot;>None</span>
    limit_for_header = <span class=&quot;num&quot;>None</span>
    <span class=&quot;kw&quot;>for</span> lim <span class=&quot;kw&quot;>in</span> limits:
        limit_scope = lim.scope <span class=&quot;kw&quot;>or</span> endpoint
        <span class=&quot;kw&quot;>if</span> lim.is_exempt(request): <span class=&quot;kw&quot;>continue</span>
        <span class=&quot;kw&quot;>if</span> lim.methods <span class=&quot;kw&quot;>is not</span> <span class=&quot;num&quot;>None</span> <span class=&quot;kw&quot;>and</span> request.method.lower() <span class=&quot;kw&quot;>not in</span> lim.methods: <span class=&quot;kw&quot;>continue</span>
        <span class=&quot;kw&quot;>if</span> lim.per_method:
            limit_scope += <span class=&quot;str&quot;>&quot;:%s&quot;</span> % request.method

        limit_key = lim.key_func(request) <span class=&quot;kw&quot;>if</span> <span class=&quot;str&quot;>&quot;request&quot;</span> <span class=&quot;kw&quot;>in</span> inspect.signature(lim.key_func).parameters <span class=&quot;kw&quot;>else</span> lim.key_func()
        args = [limit_key, limit_scope]
        <span class=&quot;kw&quot;>if</span> all(args):
            <span class=&quot;kw&quot;>if</span> self._key_prefix: args = [self._key_prefix] + args
            <span class=&quot;com&quot;># Track the SMALLEST limit -- this is what gets reported in headers.</span>
            <span class=&quot;kw&quot;>if not</span> limit_for_header <span class=&quot;kw&quot;>or</span> lim.limit &amp;lt; limit_for_header[<span class=&quot;num&quot;>0</span>]:
                limit_for_header = (lim.limit, args)

            cost = lim.cost(request) <span class=&quot;kw&quot;>if</span> callable(lim.cost) <span class=&quot;kw&quot;>else</span> lim.cost
            <span class=&quot;hi&quot;><span class=&quot;kw&quot;>if not</span> self.limiter.hit(lim.limit, *args, cost=cost):  <span class=&quot;com&quot;># &amp;lt;-- the actual check</span></span>
                failed_limit = lim
                limit_for_header = (lim.limit, args)
                <span class=&quot;kw&quot;>break</span>                                          <span class=&quot;com&quot;># stop -- don't decrement remaining limits</span>

    request.state.view_rate_limit = limit_for_header        <span class=&quot;com&quot;># picked up by header injection later</span>
    <span class=&quot;kw&quot;>if</span> failed_limit:
        <span class=&quot;kw&quot;>raise</span> RateLimitExceeded(failed_limit)</code></pre>
</div>

<div class=&quot;snippet&quot;>
<div class=&quot;snippet-title&quot;>3. Decorator entry point: <code class=&quot;inline&quot;>__limit_decorator</code> wrapper <span class=&quot;file-tag&quot;>extension.py</span></div>
<p class=&quot;annot&quot;>The <code class=&quot;inline&quot;>@limiter.limit(&quot;5/minute&quot;)</code> decorator registers the limit into <code class=&quot;inline&quot;>_route_limits</code> / <code class=&quot;inline&quot;>_dynamic_route_limits</code> at decoration time, then returns a wrapper. The wrapper finds the <code class=&quot;inline&quot;>Request</code> in args/kwargs, runs the check (passing <code class=&quot;inline&quot;>in_middleware=False</code>), runs the handler, injects headers. Note the <code class=&quot;inline&quot;>_rate_limiting_complete</code> flag -- this is the <em>only</em> guard against double-checking when middleware also fires.</p>
<pre><code><span class=&quot;kw&quot;>async def</span> async_wrapper(*args, **kwargs):
    <span class=&quot;kw&quot;>if</span> self.enabled:
        request = kwargs.get(<span class=&quot;str&quot;>&quot;request&quot;</span>, args[idx] <span class=&quot;kw&quot;>if</span> args <span class=&quot;kw&quot;>else</span> <span class=&quot;num&quot;>None</span>)
        <span class=&quot;kw&quot;>if not</span> isinstance(request, Request):
            <span class=&quot;kw&quot;>raise</span> Exception(<span class=&quot;str&quot;>&quot;parameter `request` must be an instance of starlette.requests.Request&quot;</span>)

        <span class=&quot;hi&quot;><span class=&quot;kw&quot;>if</span> self._auto_check <span class=&quot;kw&quot;>and not</span> getattr(request.state, <span class=&quot;str&quot;>&quot;_rate_limiting_complete&quot;</span>, <span class=&quot;num&quot;>False</span>):</span>
            self._check_request_limit(request, func, <span class=&quot;num&quot;>False</span>)        <span class=&quot;com&quot;># in_middleware=False</span>
            request.state._rate_limiting_complete = <span class=&quot;num&quot;>True</span>          <span class=&quot;com&quot;># &amp;lt;-- only the decorator sets this</span>

    response = <span class=&quot;kw&quot;>await</span> func(*args, **kwargs)

    <span class=&quot;kw&quot;>if</span> self.enabled:
        <span class=&quot;kw&quot;>if not</span> isinstance(response, Response):
            self._inject_headers(kwargs.get(<span class=&quot;str&quot;>&quot;response&quot;</span>), request.state.view_rate_limit)
        <span class=&quot;kw&quot;>else</span>:
            self._inject_headers(response, request.state.view_rate_limit)
    <span class=&quot;kw&quot;>return</span> response</code></pre>
</div>

<div class=&quot;snippet&quot;>
<div class=&quot;snippet-title&quot;>4. Middleware entry point <span class=&quot;file-tag&quot;>middleware.py</span></div>
<p class=&quot;annot&quot;>The middleware path is what handles <em>application limits</em> and <em>default limits</em> for routes that don't have their own decorator. <code class=&quot;inline&quot;>_should_exempt</code> deliberately skips routes that already have a static decorator-defined limit -- &quot;the decorator handles it.&quot; After the handler returns, the middleware injects headers from <code class=&quot;inline&quot;>request.state.view_rate_limit</code> (set by <code class=&quot;inline&quot;>__evaluate_limits</code>).</p>
<pre><code><span class=&quot;kw&quot;>class</span> SlowAPIMiddleware(BaseHTTPMiddleware):
    <span class=&quot;kw&quot;>async def</span> dispatch(self, request, call_next):
        app, limiter = request.app, request.app.state.limiter
        <span class=&quot;kw&quot;>if not</span> limiter.enabled:
            <span class=&quot;kw&quot;>return await</span> call_next(request)

        handler = _find_route_handler(app.routes, request.scope)
        <span class=&quot;kw&quot;>if</span> _should_exempt(limiter, handler):                <span class=&quot;com&quot;># handler missing,</span>
            <span class=&quot;kw&quot;>return await</span> call_next(request)               <span class=&quot;com&quot;># in _exempt_routes,</span>
                                                            <span class=&quot;com&quot;># OR in _route_limits (decorator owns it)</span>
        error_response, should_inject_headers = sync_check_limits(limiter, request, handler, app)
        <span class=&quot;kw&quot;>if</span> error_response <span class=&quot;kw&quot;>is not</span> <span class=&quot;num&quot;>None</span>:
            <span class=&quot;kw&quot;>return</span> error_response                          <span class=&quot;com&quot;># 429 path</span>

        response = <span class=&quot;kw&quot;>await</span> call_next(request)
        <span class=&quot;kw&quot;>if</span> should_inject_headers:
            response = limiter._inject_headers(response, request.state.view_rate_limit)
        <span class=&quot;kw&quot;>return</span> response

<span class=&quot;kw&quot;>def</span> _should_exempt(limiter, handler):
    <span class=&quot;kw&quot;>if</span> handler <span class=&quot;kw&quot;>is</span> <span class=&quot;num&quot;>None</span>: <span class=&quot;kw&quot;>return</span> <span class=&quot;num&quot;>True</span>
    name = _get_route_name(handler)
    <span class=&quot;kw&quot;>if</span> name <span class=&quot;kw&quot;>in</span> limiter._exempt_routes: <span class=&quot;kw&quot;>return</span> <span class=&quot;num&quot;>True</span>
    <span class=&quot;hi&quot;><span class=&quot;kw&quot;>if</span> name <span class=&quot;kw&quot;>in</span> limiter._route_limits: <span class=&quot;kw&quot;>return</span> <span class=&quot;num&quot;>True</span>     <span class=&quot;com&quot;># static decorator limits only -- NOT dynamic ones</span></span>
    <span class=&quot;kw&quot;>return</span> <span class=&quot;num&quot;>False</span></code></pre>
</div>

<h2>Where the three buckets of limits come from</h2>
<table>
  <thead><tr><th>Bucket</th><th>Set by</th><th>Applied when</th><th>Stored in</th></tr></thead>
  <tbody>
    <tr><td><strong>application_limits</strong></td><td><code class=&quot;inline&quot;>Limiter(application_limits=[...])</code> or <code class=&quot;inline&quot;>RATELIMIT_APPLICATION</code></td><td>Middleware path only (<code class=&quot;inline&quot;>in_middleware=True</code>). One shared bucket across the whole app, scope=<code class=&quot;inline&quot;>&quot;global&quot;</code>.</td><td><code class=&quot;inline&quot;>_application_limits</code></td></tr>
    <tr><td><strong>route limits</strong></td><td><code class=&quot;inline&quot;>@limiter.limit(...)</code> / <code class=&quot;inline&quot;>@limiter.shared_limit(...)</code></td><td>Decorator path only (<code class=&quot;inline&quot;>in_middleware=False</code>). Static strings live in <code class=&quot;inline&quot;>_route_limits</code>; callables live in <code class=&quot;inline&quot;>_dynamic_route_limits</code>.</td><td>both above</td></tr>
    <tr><td><strong>default_limits</strong></td><td><code class=&quot;inline&quot;>Limiter(default_limits=[...])</code> or <code class=&quot;inline&quot;>RATELIMIT_DEFAULT</code></td><td>Either path -- but only if the route has no decorator-applied limits, OR every route limit was declared with <code class=&quot;inline&quot;>override_defaults=False</code>.</td><td><code class=&quot;inline&quot;>_default_limits</code></td></tr>
  </tbody>
</table>

<h2>Gotchas</h2>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 1 -- The flag that prevents double-checking is set by only one side</div>
<p>The decorator wrapper sets <code class=&quot;inline&quot;>request.state._rate_limiting_complete = True</code> after running its check. The middleware <em>reads</em> that flag (in <code class=&quot;inline&quot;>_check_limits</code>) but never sets it. In practice this is fine because <code class=&quot;inline&quot;>_should_exempt</code> tells the middleware to skip routes whose name is in <code class=&quot;inline&quot;>_route_limits</code> -- but that check uses <strong>only static decorator limits</strong>. If your route has a <em>dynamic</em> limit (callable passed to <code class=&quot;inline&quot;>@limiter.limit</code>), the middleware will <em>not</em> exempt it, and you'll evaluate application/default limits in the middleware <em>and</em> the dynamic limit in the decorator on the same request.</p>
</div>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 2 -- <code class=&quot;inline&quot;>override_defaults</code> behaves the opposite of how it reads</div>
<p>It defaults to <code class=&quot;inline&quot;>True</code>, meaning &quot;if I have a route limit, skip the defaults.&quot; The merge logic is:</p>
<pre><code>combined_defaults = all(<span class=&quot;kw&quot;>not</span> l.override_defaults <span class=&quot;kw&quot;>for</span> l <span class=&quot;kw&quot;>in</span> route_limits)
<span class=&quot;kw&quot;>if not</span> route_limits <span class=&quot;kw&quot;>or</span> combined_defaults:
    all_limits += default_limits</code></pre>
<p>So defaults are added only when <em>every</em> route limit explicitly opts in by setting <code class=&quot;inline&quot;>override_defaults=False</code>. A single <code class=&quot;inline&quot;>override_defaults=True</code> limit on the route silences <em>all</em> defaults, even other route limits' opt-ins.</p>
</div>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 3 -- Headers report the <em>smallest</em> limit, not the one you might expect</div>
<p>If a route is decorated with <code class=&quot;inline&quot;>&quot;5/minute;100/hour&quot;</code>, <code class=&quot;inline&quot;>__evaluate_limits</code> tracks <code class=&quot;inline&quot;>limit_for_header</code> as the limit with the smallest <code class=&quot;inline&quot;>amount</code> seen so far. <code class=&quot;inline&quot;>X-RateLimit-Limit</code> / <code class=&quot;inline&quot;>Remaining</code> / <code class=&quot;inline&quot;>Reset</code> always reflect <em>that</em> limit. Clients reading those headers won't see the hourly bucket at all unless it becomes the binding constraint at the moment the response is built. Also: when a limit fails, the <em>failing</em> limit overwrites <code class=&quot;inline&quot;>limit_for_header</code> regardless of size, so the headers on a 429 response always describe the limit that just rejected you.</p>
</div>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 4 -- <code class=&quot;inline&quot;>hit()</code> short-circuits on the first failure</div>
<p>The loop in <code class=&quot;inline&quot;>__evaluate_limits</code> calls <code class=&quot;inline&quot;>self.limiter.hit(...)</code> for each limit in order, and <code class=&quot;inline&quot;>break</code>s on the first <code class=&quot;inline&quot;>False</code>. <code class=&quot;inline&quot;>hit</code> increments the counter as a side effect, so a 429 response means earlier limits in the list <em>were</em> incremented but later ones <em>were not</em>. If you depend on multiple counters being kept in lock-step (e.g., to derive analytics from <code class=&quot;inline&quot;>Remaining</code>), they will drift the moment any single bucket overflows. Order matters: limits are iterated in the order they were registered.</p>
</div>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 5 -- <code class=&quot;inline&quot;>SlowAPIMiddleware</code> downgrades async exception handlers</div>
<p>The classic middleware uses <code class=&quot;inline&quot;>sync_check_limits</code>, which contains:</p>
<pre><code><span class=&quot;kw&quot;>if</span> inspect.iscoroutinefunction(exception_handler):
    exception_handler = _rate_limit_exceeded_handler</code></pre>
<p>If you registered a custom <em>async</em> handler for <code class=&quot;inline&quot;>RateLimitExceeded</code> on the app, the middleware silently swaps it out for the default sync handler. Custom 429 responses, audit hooks, etc. won't run for requests that hit through this path. Use <code class=&quot;inline&quot;>SlowAPIASGIMiddleware</code> (which uses <code class=&quot;inline&quot;>async_check_limits</code>) if you need the async handler to fire.</p>
</div>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 6 -- The &quot;fallback to in-memory&quot; is best-effort and self-flips on any header error</div>
<p><code class=&quot;inline&quot;>_inject_headers</code> wraps the <code class=&quot;inline&quot;>window_stats</code> call in a bare <code class=&quot;inline&quot;>except:</code>. Any exception there -- not just storage timeouts -- causes <code class=&quot;inline&quot;>self._storage_dead = True</code> and a recursive call against the in-memory fallback. Until the next storage probe (<code class=&quot;inline&quot;>__should_check_backend</code>, exponentially backed off), every request goes to the local memory store. In a multi-worker deployment this means each worker silently gets its own private counter, and your effective rate limit is multiplied by worker count. The probe will eventually re-test the real backend and recover -- but the window of inconsistency depends on <code class=&quot;inline&quot;>MAX_BACKEND_CHECKS</code> (5) and <code class=&quot;inline&quot;>2**check_count</code> seconds.</p>
</div>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 7 -- Two latent bugs in the codebase</div>
<ul style=&quot;margin: 4px 0 0;&quot;>
  <li><code class=&quot;inline&quot;>slowapi_startup</code> references a bare name <code class=&quot;inline&quot;>app</code> with no enclosing scope (see the <code class=&quot;inline&quot;># type: ignore</code>s). Calling it raises <code class=&quot;inline&quot;>NameError</code>. The intended wiring is <code class=&quot;inline&quot;>app.state.limiter = limiter</code> done by the user.</li>
  <li><code class=&quot;inline&quot;>Limit.scope</code> calls <code class=&quot;inline&quot;>self.__scope(request.endpoint)</code> when the scope is callable -- but <code class=&quot;inline&quot;>request</code> is undefined in that property and <code class=&quot;inline&quot;>Request</code> objects don't have <code class=&quot;inline&quot;>.endpoint</code> in Starlette anyway. The in-source <code class=&quot;inline&quot;>FIXME</code> confirms this path is broken; callable scopes will raise.</li>
</ul>
</div>

<div class=&quot;gotcha&quot;>
<div class=&quot;label&quot;>Gotcha 8 -- The <code class=&quot;inline&quot;>key_style</code> choice silently changes which routes share buckets</div>
<p>With <code class=&quot;inline&quot;>key_style=&quot;url&quot;</code> (the default) the bucket scope is the literal request path, so <code class=&quot;inline&quot;>/users/1</code> and <code class=&quot;inline&quot;>/users/2</code> have <em>separate</em> rate-limit buckets even though they hit the same handler. With <code class=&quot;inline&quot;>key_style=&quot;endpoint&quot;</code>, both share a bucket. If you have parameterized routes and you set up a per-endpoint quota expecting &quot;10 calls per minute against this handler,&quot; the default URL keying gives you &quot;10 per minute per distinct path you happen to call&quot; instead. Pick deliberately.</p>
</div>

</div>
</body>
</html>
" height="700" width="100%" style="border:1px solid #d0d7de;border-radius:6px;margin:1.5em 0"></iframe>

*The HTML rate-limiter explainer. Mermaid flowchart via CDN, callout boxes for the 8 gotchas, syntax-highlighted code via CSS classes. Looks the part.*

<iframe srcdoc="<!doctype html>
<html lang=&quot;en&quot;>
<head>
<meta charset=&quot;utf-8&quot;>
<title>Rendered markdown</title>
<style>
  body { font: 15px/1.6 -apple-system, BlinkMacSystemFont, &quot;Segoe UI&quot;, Roboto, sans-serif;
         color: #1f2328; background: #ffffff; max-width: 760px; margin: 0 auto; padding: 24px; }
  h1 { font-size: 24px; margin-top: 0; border-bottom: 1px solid #d0d7de; padding-bottom: 8px; }
  h2 { font-size: 20px; margin-top: 28px; border-bottom: 1px solid #d0d7de; padding-bottom: 6px; }
  h3 { font-size: 16px; margin-top: 22px; }
  code { background: #f6f8fa; padding: 1px 5px; border-radius: 4px; font-size: 13px;
         font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace; }
  pre { background: #f6f8fa; padding: 12px 14px; border-radius: 6px; overflow-x: auto;
        font-family: ui-monospace, SFMono-Regular, Menlo, Consolas, monospace; font-size: 13px;
        line-height: 1.45; }
  pre code { background: transparent; padding: 0; font-size: inherit; }
  blockquote { border-left: 3px solid #d0d7de; padding-left: 12px; color: #59636e; margin: 12px 0; }
  table { border-collapse: collapse; margin: 12px 0; font-size: 14px; }
  th, td { border: 1px solid #d0d7de; padding: 6px 12px; text-align: left; vertical-align: top; }
  th { background: #f6f8fa; }
  hr { border: none; border-top: 1px solid #d0d7de; margin: 28px 0; }
  a { color: #0969da; }
</style>
</head>
<body>
<h1>slowapi rate limiter, explained</h1>
<p>slowapi rate-limits requests to a Starlette/FastAPI app. It is a thin coordinator on top of the <code>limits</code> library: <code>limits</code> owns the storage (memory/Redis/...) and the strategy (fixed-window/...). slowapi owns <em>which</em> limits apply to <em>this</em> request, <em>what key</em> identifies the caller, and <em>what headers</em> go on the response.</p>
<p>There are two ways limits get registered, and two ways they get checked. Keeping those four straight is most of understanding the code.</p>
<h2>How limits get registered</h2>
<table>
<thead>
<tr>
<th>Source</th>
<th>Registered into</th>
<th>Applied by</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>Limiter(default_limits=[...])</code></td>
<td><code>_default_limits</code></td>
<td>both paths, when route has no overriding limit</td>
</tr>
<tr>
<td><code>Limiter(application_limits=[...])</code></td>
<td><code>_application_limits</code></td>
<td>middleware path only</td>
</tr>
<tr>
<td><code>@limiter.limit(&quot;5/minute&quot;)</code> (static)</td>
<td><code>_route_limits[name]</code></td>
<td>decorator wrapper</td>
</tr>
<tr>
<td><code>@limiter.limit(callable)</code> (dynamic)</td>
<td><code>_dynamic_route_limits[name]</code></td>
<td>decorator wrapper</td>
</tr>
<tr>
<td><code>@limiter.exempt</code></td>
<td><code>_exempt_routes</code></td>
<td>both paths skip</td>
</tr>
</tbody>
</table>
<p>A &quot;limit&quot; is parsed by <code>limits.parse_many(&quot;5/minute;100/hour&quot;)</code> into one or more <code>RateLimitItem</code> objects, each wrapped in a <code>Limit</code> (wrappers.py). <code>LimitGroup</code> is iterable and yields one <code>Limit</code> per parsed item.</p>
<h2>How a request is checked</h2>
<pre><code class=&quot;language-mermaid&quot;>flowchart TD
    A[Request arrives] --&amp;gt; B{Middleware&amp;lt;br/&amp;gt;installed?}
    B --&amp;gt;|Yes| C[SlowAPIMiddleware.dispatch]
    B --&amp;gt;|No| K

    C --&amp;gt; D{_should_exempt?&amp;lt;br/&amp;gt;handler in&amp;lt;br/&amp;gt;_exempt_routes or&amp;lt;br/&amp;gt;_route_limits}
    D --&amp;gt;|Yes, skip&amp;lt;br/&amp;gt;middleware checks| K[call route handler]
    D --&amp;gt;|No| E[_check_request_limit&amp;lt;br/&amp;gt;in_middleware=True]
    E --&amp;gt; F[Builds: application_limits&amp;lt;br/&amp;gt;+ defaults if applicable]
    F --&amp;gt; G[__evaluate_limits]
    G --&amp;gt;|hit OK| H[continue to handler]
    G --&amp;gt;|hit fails| X[raise RateLimitExceeded&amp;lt;br/&amp;gt;→ 429 JSON response]
    H --&amp;gt; K

    K --&amp;gt; L{Route has&amp;lt;br/&amp;gt;@limiter.limit?}
    L --&amp;gt;|No| Z[response]
    L --&amp;gt;|Yes| M[decorator's sync/async_wrapper]
    M --&amp;gt; N{_rate_limiting_complete&amp;lt;br/&amp;gt;already True?}
    N --&amp;gt;|Yes, already checked| O[run handler]
    N --&amp;gt;|No| P[_check_request_limit&amp;lt;br/&amp;gt;in_middleware=False]
    P --&amp;gt; Q[Builds: route_limits&amp;lt;br/&amp;gt;+ dynamic_limits&amp;lt;br/&amp;gt;+ defaults if applicable]
    Q --&amp;gt; R[__evaluate_limits]
    R --&amp;gt;|hit OK| S[set _rate_limiting_complete=True]
    R --&amp;gt;|hit fails| X
    S --&amp;gt; O
    O --&amp;gt; T[_inject_headers&amp;lt;br/&amp;gt;X-RateLimit-* + Retry-After]
    T --&amp;gt; Z

    style X fill:#fdd
    style G fill:#ffd
    style R fill:#ffd
</code></pre>
<p>The decorator and middleware paths are <strong>independent</strong>. The middleware checks if a route has a registered <code>@limiter.limit</code> — if so, it skips its own check and lets the decorator handle it (<code>_should_exempt</code> returns True for routes in <code>_route_limits</code>). Application-wide limits therefore only fire on routes that do <em>not</em> carry a decorator, unless you also wrap them with the decorator separately. This is the most surprising design choice in the library.</p>
<h2>Key code, annotated</h2>
<h3>1. The decorator wrapper — where rate-limiting actually attaches to a route</h3>
<p><code>extension.py</code>, inside <code>__limit_decorator</code>:</p>
<pre><code class=&quot;language-python&quot;>@functools.wraps(func)
async def async_wrapper(*args, **kwargs):
    if self.enabled:
        request = kwargs.get(&amp;quot;request&amp;quot;, args[idx] if args else None)
        # Idempotency guard: if the middleware already checked, don't double-hit storage.
        if self._auto_check and not getattr(
            request.state, &amp;quot;_rate_limiting_complete&amp;quot;, False
        ):
            self._check_request_limit(request, func, False)  # in_middleware=False
            request.state._rate_limiting_complete = True
    response = await func(*args, **kwargs)
    if self.enabled:
        # Even if the check was skipped (already complete), headers still get injected
        # using request.state.view_rate_limit set during evaluation.
        self._inject_headers(response, request.state.view_rate_limit)
    return response
</code></pre>
<p>The <code>_rate_limiting_complete</code> flag is only ever set by the decorator wrapper — never by the middleware. So middleware → decorator double-checking is prevented by the middleware's <code>_should_exempt</code> returning early for decorated routes, <em>not</em> by the flag. The flag exists for the rarer case of nested decorators / repeated dispatch.</p>
<h3>2. Building the limit list — when do defaults apply?</h3>
<p><code>extension.py</code>, inside <code>_check_request_limit</code>:</p>
<pre><code class=&quot;language-python&quot;>route_limits: List[Limit] = limits + dynamic_limits
all_limits = (
    list(itertools.chain(*self._application_limits))
    if in_middleware
    else []
)
all_limits += route_limits
combined_defaults = all(
    not limit.override_defaults for limit in route_limits
)
if (
    not route_limits
    and not (in_middleware and endpoint_func_name in self.__marked_for_limiting)
    or combined_defaults
):
    all_limits += list(itertools.chain(*self._default_limits))
</code></pre>
<p>Read the boolean carefully — Python parses it as <code>(A and not B) or C</code>:</p>
<ul>
<li><strong>A</strong>: this route has no route-level limits, AND</li>
<li><strong>B</strong> (negated): we're not in the middleware-skipping-a-decorated-route case, OR</li>
<li><strong>C</strong>: every route-level limit was registered with <code>override_defaults=False</code>.</li>
</ul>
<p>So defaults stack with route limits <em>only</em> when you opt out of overriding via <code>override_defaults=False</code> (note: the decorator default is <code>True</code>, i.e. route limits replace defaults by default). This is the second-most-surprising design choice.</p>
<h3>3. The hit loop — where requests are actually counted</h3>
<p><code>extension.py</code>, <code>__evaluate_limits</code>:</p>
<pre><code class=&quot;language-python&quot;>def __evaluate_limits(self, request, endpoint, limits):
    failed_limit = None
    limit_for_header = None
    for lim in limits:
        limit_scope = lim.scope or endpoint
        if lim.is_exempt(request): continue
        if lim.methods is not None and request.method.lower() not in lim.methods: continue
        if lim.per_method:
            limit_scope += &amp;quot;:%s&amp;quot; % request.method

        # key_func may or may not accept `request` — sniffed via inspect.signature
        if &amp;quot;request&amp;quot; in inspect.signature(lim.key_func).parameters.keys():
            limit_key = lim.key_func(request)
        else:
            limit_key = lim.key_func()

        args = [limit_key, limit_scope]
        if all(args):  # silently skip if either is empty/falsy
            if self._key_prefix:
                args = [self._key_prefix] + args
            # Track the SMALLEST limit (most restrictive) for headers.
            if not limit_for_header or lim.limit &amp;lt; limit_for_header[0]:
                limit_for_header = (lim.limit, args)

            cost = lim.cost(request) if callable(lim.cost) else lim.cost
            if not self.limiter.hit(lim.limit, *args, cost=cost):
                # First failure wins — break, do NOT hit subsequent limits.
                failed_limit = lim
                limit_for_header = (lim.limit, args)
                break
        else:
            self.logger.error(&amp;quot;Skipping limit: %s. Empty value found in parameters.&amp;quot;, lim.limit)
            continue
    request.state.view_rate_limit = limit_for_header
    if failed_limit:
        raise RateLimitExceeded(failed_limit)
</code></pre>
<p>Two points worth burning in:</p>
<ol>
<li>The header-target is the <strong>smallest</strong> limit, by <code>RateLimitItem.__lt__</code> (smaller window/amount = more restrictive). This is what the client sees in <code>X-RateLimit-*</code>.</li>
<li>On failure, the loop <strong>breaks</strong> — slowapi reports the first limit that failed and never increments counters for limits later in the list. This matters when a route has multiple stacked limits (&quot;5/sec;100/min;1000/hr&quot;); only the smallest-window-that-fails counts the request.</li>
</ol>
<h3>4. Header injection — the contract with the client</h3>
<p><code>extension.py</code>, <code>_inject_headers</code> (asgi variant is identical, just operating on <code>MutableHeaders</code>):</p>
<pre><code class=&quot;language-python&quot;>window_stats = self.limiter.get_window_stats(current_limit[0], *current_limit[1])
reset_in = 1 + window_stats[0]   # window_stats[0] is the unix ts when the window resets
response.headers.append(self._header_mapping[HEADERS.LIMIT], str(current_limit[0].amount))
response.headers.append(self._header_mapping[HEADERS.REMAINING], str(window_stats[1]))
response.headers.append(self._header_mapping[HEADERS.RESET], str(reset_in))

existing_retry_after_header = response.headers.get(&amp;quot;Retry-After&amp;quot;)
if existing_retry_after_header is not None:
    reset_in = max(self._determine_retry_time(existing_retry_after_header), reset_in)
response.headers[self._header_mapping[HEADERS.RETRY_AFTER]] = (
    formatdate(reset_in) if self._retry_after == &amp;quot;http-date&amp;quot;
    else str(int(reset_in - time.time()))
)
</code></pre>
<p><code>Retry-After</code> defaults to delta-seconds (<code>reset_in - now</code>); set <code>retry_after=&quot;http-date&quot;</code> to emit RFC2616 dates instead. Headers are only injected if <code>headers_enabled=True</code> <em>and</em> a limit was evaluated for this request — otherwise the response is silently un-annotated.</p>
<h2>Gotchas</h2>
<ol>
<li>
<p><strong>Application limits only fire on un-decorated routes.</strong> <code>SlowAPIMiddleware</code> short-circuits via <code>_should_exempt</code> for any route present in <code>_route_limits</code>. If you set <code>application_limits=[&quot;1000/hour&quot;]</code> and decorate every endpoint, the application limit never runs. Either keep the application limit's routes un-decorated, or stack the limit explicitly.</p>
</li>
<li>
<p><strong><code>override_defaults=True</code> is the decorator default.</strong> Adding <code>@limiter.limit(&quot;5/minute&quot;)</code> <em>replaces</em> your <code>default_limits</code>, it does not stack. Pass <code>override_defaults=False</code> to stack. The variable <code>combined_defaults</code> in <code>_check_request_limit</code> reads as if it computes a union — it does not; it just decides whether to <em>append</em> defaults.</p>
</li>
<li>
<p><strong>Empty <code>key_func()</code> or <code>scope</code> silently skips a limit.</strong> The <code>if all(args):</code> guard means a buggy key function that returns <code>&quot;&quot;</code> or <code>None</code> (e.g., user not authenticated, header missing) results in a request that is <em>not</em> rate-limited at all, with only an error log. This is a common bypass vector — always make <code>key_func</code> return a sentinel like <code>&quot;anonymous&quot;</code> rather than an empty string.</p>
</li>
<li>
<p><strong>Multiple stacked limits: only the first-failing one counts the hit.</strong> A request that would breach <code>5/sec</code> and <code>100/min</code> simultaneously increments the per-second counter and stops; the per-minute counter is untouched for that request. So aggregate &quot;minute&quot; usage will under-report whenever the second-window limit is being hit.</p>
</li>
<li>
<p><strong>The headers report the <em>smallest</em> limit, not the first.</strong> If you stack <code>1000/hr;5/sec</code>, <code>X-RateLimit-Limit</code> will be <code>5</code>. Clients implementing backoff against a single header pair will see the tightest limit, which is usually what you want — but be aware your &quot;1000/hr&quot; budget is invisible to them.</p>
</li>
<li>
<p><strong><code>Limit.scope</code> is broken for callable scopes</strong> (<code>wrappers.py</code>). The property body references an undefined <code>request.endpoint</code> (a Flask-ism leftover, with a <code>FIXME</code> in the source). Passing a callable as <code>scope=</code> to <code>shared_limit</code> will raise <code>NameError</code> at evaluation time. Use string scopes only.</p>
</li>
<li>
<p><strong><code>scope=</code> on <code>@limiter.limit()</code> is a no-op.</strong> <code>__limit_decorator</code> does <code>_scope = scope if shared else None</code>, and <code>limit()</code> calls it without <code>shared=True</code>. Only <code>shared_limit()</code> actually uses scope.</p>
</li>
<li>
<p><strong>In-memory fallback hides backend outages.</strong> When <code>_storage_dead = True</code> and <code>_in_memory_fallback_enabled</code>, slowapi uses local memory until the backend recovers. In a multi-process deployment each worker maintains its own counter while the backend is down, so the effective limit during outages is <code>N_workers × configured_limit</code>. Health is rechecked with exponential backoff (<code>pow(2, count)</code>), so recovery can take up to ~32 seconds after the backend returns.</p>
</li>
<li>
<p><strong><code>slowapi_startup</code> is dead code.</strong> It references a free variable <code>app</code> and would <code>NameError</code> if called. Don't wire it up; set <code>app.state.limiter = limiter</code> and register the <code>RateLimitExceeded</code> exception handler manually (or use the middleware, which reads <code>app.state.limiter</code>).</p>
</li>
<li>
<p><strong><code>SlowAPIMiddleware</code> (BaseHTTPMiddleware) only handles sync exception handlers</strong>; if your custom <code>RateLimitExceeded</code> handler is async it silently falls back to the default <code>_rate_limit_exceeded_handler</code>. Use <code>SlowAPIASGIMiddleware</code> if you need async exception handlers.</p>
</li>
<li>
<p><strong><code>key_func</code> signature is sniffed, not specified.</strong> <code>__evaluate_limits</code> introspects <code>key_func.parameters</code> and calls <code>key_func(request)</code> if <code>&quot;request&quot;</code> is in the signature, else <code>key_func()</code>. A decorator like <code>@functools.wraps</code> that hides parameters, or a <code>lambda *a: ...</code>, will land on the wrong branch.</p>
</li>
<li>
<p><strong><code>request.state.view_rate_limit</code> is the contract between check and header injection.</strong> If you wrap the response in custom middleware after slowapi without preserving <code>request.state</code>, headers won't be injected. The exception handler <code>_rate_limit_exceeded_handler</code> also reads from it — clearing it breaks 429 headers.</p>
</li>
</ol></body></html>" height="700" width="100%" style="border:1px solid #d0d7de;border-radius:6px;margin:1.5em 0"></iframe>

*The same explainer, in markdown. Twelve gotchas instead of eight. Mermaid flowchart in a fenced block, fenced code blocks letting the viewer handle syntax highlighting, no CSS chrome. Cheaper. And, to my eye, the better artifact.*

The HTML version's distinctive features are *CSS chrome*: a stylesheet for syntax-highlighting tokens (`<span class="kw">`, `<span class="str">`), background panels with `border-radius` for code blocks, color-coded callout boxes for gotchas. The markdown version uses fenced code blocks which most modern viewers syntax-highlight automatically.

I counted the gotchas in both files. HTML had 8. Markdown had 12. The HTML compresses two findings into one "Two latent bugs" entry, so charitably call it 9-10. Even charitably, the markdown surfaces more distinct gotchas the HTML doesn't reach: the empty-`key_func` bypass, the `inspect.parameters` signature sniffing, the `request.state.view_rate_limit` contract between checking and header injection. The HTML caught one the markdown missed (the `key_style="url"` vs `"endpoint"` distinction), but it isn't a wash. Markdown was the more substantive artifact.

That surprised me. Going in, I'd assumed the verbose-priming effect (HTML's tag-heavy output register leading Claude into more verbose responses) would push HTML toward more comprehensive answers. It didn't here. It pushed the other way.

Think of it like two technical writers given the same source code. One hands you a clean PDF with a cover page, syntax-highlighted snippets, and tasteful colored boxes around the warnings. The other gives you a longer markdown file with no styling, but he caught four more gotchas. Which one do you want before you ship to prod?

The edit probe threw me a small surprise. I asked Claude to add a "Custom storage backends" section to each artifact and looked at the diffs. HTML's diff was 2,471 bytes over 19 lines; markdown's was *bigger* at 2,987 bytes over the same 19 lines. Same line count, but markdown's paragraph-style prose has longer lines than HTML's tag-broken structure. The "HTML diffs are noisier" claim turns out not to generalize cleanly. When the content is mostly prose, markdown isn't a saving on diff size.

Verdict, with the obvious caveat that this is one sample on one substrate: the markdown was the better artifact. More comprehensive (12 gotchas vs effectively 9-10 once you count the HTML's "two latent bugs" as separate items), cheaper, with a real Mermaid diagram in the fenced block. The intuition that "HTML wins because richer information" doesn't hold up in this case. Whether it'd hold up at K=5 generations I genuinely don't know, and I think that's the right answer to give. One sample is one sample. But the direction of the surprise is striking. I expected HTML's verbose-priming effect to push HTML toward more thorough answers; it pushed the other way here.

## "The 1M context window handles it." Does it though?

Thariq's FAQ says: *"With the 1M context window in Opus 4.7, the increased token usage is not really noticeable in the context window."*

Fair point for one-shot generation. Not so fair for re-ingestion, which is what Claude Code actually does. Feed yesterday's artifact back in as context for today's follow-up turn, over and over. Let me show what that costs.

**Cache_creation tokens (artifact-as-input-context) per re-ingestion across the three cases:**

| Use case | HTML | MD | Delta |
|---|---|---|---|
| Design exploration | 27,309 | 21,657 | +5,652 |
| PR review          | 33,892 | 25,532 | +8,360 |
| Rate limiter expl. | 32,280 | 26,379 | +5,901 |

Run the math. Project produces N artifacts. Each gets re-fed into the agent K times. The cache evicts between feeds (i.e. calls more than ~5 minutes apart, where you pay the cache_creation cost fresh). The HTML-over-markdown overhead is roughly N × K × 6,000 tokens. At N=50, K=3 that's 900,000 tokens of cache_creation, basically the whole 1M context window, spent on markup you wouldn't have if you'd used markdown. For tight call chains where the cache stays warm, subsequent reads hit `cache_read_input_tokens` (much cheaper, \$1.50/M vs \$15/M for fresh input), and the overhead amortizes to roughly the artifact's first-ingest cost.

The dollar cost differential is muted by cache pricing for tight chains, so individual calls feel cheap. The *context budget* is where the cost most reliably shows, because that's measured in tokens regardless of cache state.

This matters more if you're not on an enterprise plan with effectively unlimited tokens. Several commenters in Thariq's thread raised exactly that point: the "context budget isn't a concern" framing reads differently from inside Anthropic than it does for a developer paying retail. The math behind that pushback is real.

## Does Claude read HTML better than markdown? (Spoiler: not really.)

Several commenters on Thariq's piece pushed back on the agent side of his argument. HTML "wastes the model's cognitive space on tags, nesting, closure, styles" (@AlexMares). It has "low signal-to-noise ratio" and risks hallucination (@AkhilDevelops). Markdown is just the right format for agents to read (@AiAGiAI). And Thariq's whole pitch quietly assumes the opposite: that the agent processes HTML at least as well as it processes markdown. Otherwise his "use HTML for everything, Claude reads it back later" workflow doesn't really hold up, does it?

So I tested it. Seven specific factual questions about slowapi's rate limiter, each answerable from both the HTML and the markdown version of the same explainer artifact (shared content, different format). I asked Claude (fresh session, Opus 4.7) to answer all seven questions from each format, then scored the answers against ground truth derived from the slowapi source.

Here's what came back:

| Metric | HTML source | MD source | Ratio HTML/MD |
|---|---|---|---|
| Correct answers | 7 / 7 | 7 / 7 | equal |
| Input cache_creation tokens | 30,213 | 24,312 | 1.24x |
| Output tokens generated | 1,560 | 785 | 1.99x |
| Cost per call | $0.246 | $0.185 | 1.33x |

Both formats produced equally accurate answers to all seven questions, confirmed by two independent judges scoring blinded (neither judge knew which set came from which format). Both judges gave both sets 7/7 on factual accuracy.

On a separate axis (explanatory specificity, the "clarity and learning value" of the answer) both judges scored HTML marginally higher. Judge 1 gave markdown 6.5/7 and HTML 7.0/7. Judge 2 gave markdown 6.0/7 and HTML 6.5/7. The HTML-sourced answers consistently cited more specific code constructs (e.g. `all(not l.override_defaults for l in route_limits)`, `MAX_BACKEND_CHECKS=5`, `inspect.iscoroutinefunction`). The markdown-sourced answers were slightly tighter prose. Both judges independently described the same pattern.

So the strong commenter worry, that HTML "wastes the model's cognitive space" or "lowers signal-to-noise to the point of hallucination," didn't show up. If anything, the mild *opposite* did: HTML-sourced answers came back a hair more specific. Both blinded judges said so. But the price-quality math is bad either way: you pay 33% more per call for answers that score ~7% better on specificity and identically on accuracy. The verbose-priming story holds up at the qualitative level too. HTML doesn't spend its extra tokens on factual errors; it spends them on more specific code citations.

A few things to keep in mind before reading too much into this:

- N=1 per condition. Single shot per format. Could be variance.
- Tested only on shared content in a single artifact. I didn't try cases where the two formats actually differ in what they covered.
- Didn't test whether answer *quality* beyond factual accuracy matters for downstream agent reasoning. Two answers can both be correct and still be differently useful for the agent's next turn.

What the test does is rule out the strong version of the agent-side worry. HTML doesn't degrade accuracy on shared content. It does cost more on both directions of the interaction for the same outcome.

## The two I didn't bother instrumenting (because there's no markdown alternative)

**Checkout button prototype.** Thariq's prompt: *"...Create a HTML file with several sliders and options for me to try different options or actions. Eventually, give me a copy button..."* This is a working interactive widget. Markdown has no equivalent. The cost question is moot; there is no markdown alternative to compare against.

**Linear ticket reorderer.** Thariq's prompt: *"Make me an HTML file with each ticket as a draggable card across New / Next / Later / Cut columns..."* Drag-and-drop is an HTML+JavaScript capability. Markdown cannot do this.

Both of these are real use cases. They demonstrate the categorical-win bucket. The token cost is the price of admission, and there is no marginal-cost-vs-benefit question because there is no marginal alternative.

## Going through Thariq's claims, one by one

| Claim from the article | My verdict |
|---|---|
| HTML has higher information density than markdown | True for visual content; false for text-only content. Markdown handles prose, tables, code blocks well. |
| HTML is more readable past 100 lines | True for visual artifacts (confirmed by user reading observation); equivocal for text-with-structure (markdown's structure works well at length). |
| HTML is easier to share than markdown | Partly true (raw browsers don't render `.md` files), but understated: GitHub renders markdown natively (and renders Mermaid as real diagrams since 2022); VS Code's built-in preview renders markdown; Slack's own "mrkdwn" subset handles a useful chunk. The "you have to attach as email" framing in the article ignores that uploading to GitHub/Gist and sharing a link is a one-step workaround. Untestable here at engagement-level. |
| HTML enables two-way interaction | True categorically. This is the strongest argument in the article (categorical-win bucket). |
| "Token overhead is not really noticeable" with 1M context | False. Per-artifact overhead is 1.4-1.9x in dollars, 2-4x in tokens. Cumulative overhead approaches the 1M context limit for moderate-sized projects. |
| "2-4x slower" generation | Partially supported: measured 1.4-2.4x across the three cases. The high end of his range is reachable; the low end is below what I measured. |
| HTML diffs are noisy (acknowledged downside) | Confirmed for verbatim methodology (1.2-2.5x diff bytes). Less clear for text-heavy content where markdown's prose lines have their own size. |
| HTML is more joyful to work with | Subjective. Skipped. |
| **"HTML is strictly better even for text-focused content"** (his hardline reply) | **Not what I saw in the data.** Both text cases produced competent markdown that did the job, at 1.4-1.5x less cost than the HTML. One of them produced *more* substance than the HTML. N=1 per condition isn't enough to refute the hardline definitively, but it's plenty to say the strong version of the claim isn't warranted by what I measured. |
| *Implied* claim: HTML reads better as agent context downstream | **Mixed.** On shared-content factual retrieval (7-question test on slowapi gotchas), two blinded judges scored both formats 7/7 on accuracy. On explanatory specificity, both judges gave HTML a ~7% edge (6.5-7.0 vs 6.0-6.5). HTML costs 1.33x more per call (1.24x input + 1.99x output). HTML-sourced answers are marginally more specific (more code-citation detail) at significantly higher cost; the quality-per-dollar trade favors markdown. |

## So what should you actually do?

**HTML if** the task is fundamentally visual or interactive: design exploration, prototypes, mockups, dashboards, drag-and-drop widgets, slider-tuned UIs, throwaway editor surfaces. Or you specifically need to send it to someone who won't render markdown themselves. The token cost is what the capability costs. Pay it without thinking.

**Markdown if** the task is text with structure: PR reviews, postmortems, design docs, implementation plans, status reports, meeting notes, technical explainers. The substance lives in the words. HTML's polish costs about 40-50% more per generation, and in my one rate-limiter sample (take it for what one sample's worth) that polish came at the expense of substance, not on top of it.

**Maybe HTML on a text task if** the artifact is going to be the centerpiece of a meeting, or you need visual hierarchy to handle dense reference material (tables of tables, comparison matrices), or the reader can't comfortably read raw markdown. Otherwise, save the tokens.

The "HTML for everything" stance treats two different questions as one. Half of it is true and worth paying for. The other half is paying extra for things the markdown was already going to give you. Pick the right tool for the right job.

## If you want to try this yourself

You don't need anything fancy. Opus 4.7 (the model Thariq is pitching for) via Claude Code's non-interactive mode does the whole thing: `claude --print --model claude-opus-4-7 -- "<the prompt>"`. The prompts are scattered through this post in blockquotes; use them verbatim for the HTML-affording methodology, or rephrase to drop the format-specific cues if you want the neutral version.

For substrate, pick something real. A non-trivial recent commit from a codebase you know for the PR review case. Any reasonably-sized library you don't fully understand for the explainer case (I used [slowapi](https://github.com/laurentS/slowapi)). Anything plausible and visually rich for the design exploration.

For token counts, I used [`toks`](https://pypi.org/project/toks/), a small Anthropic-tokenizer wrapper that reports an exact token count for whichever provider you're targeting. Without it I would have been squinting at character counts and guessing.

One thing I'd do differently if I were starting over: run K=5 generations per condition. Mine were all K=1, which is why a lot of the findings (especially "markdown was more substantive on the rate-limiter") come with the N=1 caveat. If your finding is going to be load-bearing, run it five times.

## What this cost me

The whole thing. Three instrumented cases, two methodologies on one of them, three probes each, the agent-as-reader experiment, plus the two demo cases. Ran me $9.42 across 26 API calls on Opus 4.7. About 28 minutes of API time, less in wall clock once you parallelize. Cheaper than the lunch I had while it ran. Tracking?
