The HTML Parser: From Bytes to the DOM Tree (and Why One <script> Halts Parsing)
In the previous article we watched the main thread and its scheduler from a distance: we know rendering is the third step of the event loop, and that synchronous JavaScript blocks rendering.
But a more fundamental question has been hanging in the air: when the browser receives a chunk of HTML, what actually happens?
"View page source" shows you a bunch of <div>s, <p>s and text. Yet in JavaScript it becomes a DOM tree you can query with document.querySelector('.title'). Who performs that translation?
Even more counterintuitive: drop a <script> containing an infinite loop into a page, and the whole thing goes blank and freezes. How can a bit of JavaScript stop HTML from even being parsed?
In this article we'll follow the browser from the bytes on the network all the way to the DOM tree you manipulate. Once you see this chain, a lot of "truisms" (put scripts at the bottom, use defer, don't dump big JS on the first screen) stop being things to memorize.
What You'll Learn
- How a chunk of HTML goes from "bytes" to "DOM tree" step by step
- Why the HTML parser is "incremental" (parses while downloading) and "error-tolerant" (fixes broken markup)
- Why a plain
<script>pauses HTML parsing - What actually differs between
asyncanddefer - Where CSS comes in (the CSSOM), and why it "blocks rendering"
- The difference between
DOMContentLoadedandload
Prerequisites
- It helps to read the event loop article first, to understand the main thread and "rendering is step 3 of the loop":
- No compiler-theory background needed; this starts from zero
Starting from a Scenario
Suppose the server sends the browser this HTML:
<!DOCTYPE html>
<html>
<head>
<link rel="stylesheet" href="style.css">
</head>
<body>
<h1 class="title">Hello</h1>
<p>world</p>
<script src="app.js"></script>
<p>after script</p>
</body>
</html>The browser ultimately needs to end up with a tree it can manipulate:
Document
└── html
├── head
│ └── link
└── body
├── h1.title ("Hello")
├── p ("world")
├── script
└── p ("after script")Now the questions:
- What the network gave the browser was just a sequence of bytes — how did it become this tree?
- If
app.jscontains awhile(true){}, why does even<p>after script</p>fail to appear?
The answer to the second question will explain a whole family of performance behaviors we take for granted.
Core Content
1. One Pipeline: Bytes → Characters → Tokens → DOM Tree
What the browser receives as HTML is fundamentally a byte stream (a pile of 0s and 1s). It isn't text, and it certainly isn't a tree. To become the DOM, it goes through a few steps:
network → byte stream
│ ① decode (using the declared encoding, usually UTF-8)
▼
character stream <html><body>hello</body>...
│ ② tokenize
▼
token sequence {StartTag: html} {StartTag: body} {Text: "hello"} {EndTag: body}
│ ③ tree construction
▼
DOM tree (a tree of nodes)- Decode: bytes have no "character meaning" by themselves. The browser determines the encoding from hints like the BOM,
Content-Type, and<meta charset>, and turns bytes back into characters (this step is called "encoding sniffing" in the HTML standard). - Tokenize: cut the character stream into meaningful tokens — start tags, end tags, text, comments, the DOCTYPE, and so on.
- Tree construction: assemble those tokens into a tree following their nesting — a start tag hangs a node below, an end tag moves back up.
The key point: these three steps are not a three-stage process where everything is decoded, then everything is tokenized, then everything is built. They're one pipeline — a little bytes come in, and it advances a little. Understanding this is the key to the next section on "incremental" parsing.
2. Incremental Parsing: Parse While You Download
The network doesn't hand the whole HTML to the browser at once; it arrives in chunks. So the browser does something bold and clever: the parser doesn't wait for the HTML to arrive complete — it parses as it comes.
chunk 1 received: <div><p>hello
→ the parser can already create the div and p nodes and put in some text
chunk 2 received: </p></div><span>...
→ keep on buildingThis is incremental parsing, also called streaming parsing. It has a few direct consequences:
- Faster first paint: the browser doesn't need the full 1 MB of HTML before it can start rendering the beginning.
- You can touch the DOM mid-parse: that's why you can attach events to
document.bodyand script can modify nodes that have already appeared.
Think about the alternative: if the browser insisted on the full download before building the tree, a large page's first paint would wait a long time. Incremental parsing is one of the foundations of the web's feel.
3. Error-Tolerant Parsing: HTML Is Not "Strict Mode"
If you've written XML, you know how strict it is: one unclosed tag or misplaced nesting and it errors out and refuses to work.
HTML is the opposite — its parsing rules are error-tolerant. Faced with broken, scrambled markup, the browser doesn't error; it "repairs" it according to an established set of rules and renders the page as best it can. A few examples you'll run into:
you wrote: <p>first<p>second
browser reads: <p>first</p><p>second</p> ← closes tags for you
you wrote: <table><div>misfiled</div></table>
browser reads: move the div outside the table (foster parenting) ← builds structure for youWhy is HTML so forgiving? Because the web must be backward compatible. Historically, so many pages were written incorrectly that if browsers crashed on the first mistake like XML does, those pages would have become unusable. So the standards authors chose to "render as best as possible."
This leads to an important fact: the source you write and the DOM tree you get are not necessarily the same shape. The parser may complete, move, or even discard some of your nodes. What DevTools' Elements panel shows is the parsed DOM, not your source.
4. The Crucial Part: Why <script> Pauses Parsing
Now we reach the most important section.
When the parser encounters a plain <script> (with neither async nor defer) in the HTML, it does something that looks brutal:
pause HTML parsing
→ download and immediately execute this JavaScript
→ when it finishes, resume parsingThis is what parser-blocking means. Why stop? Not because the browser is lazy, but because a script might change the document in turn:
- The script might call
document.write(...)to insert HTML into the document; - The script might query the DOM (
document.querySelector(...)) and expects to see "the fully parsed DOM up to this point."
If the browser kept parsing while the script freely changed things, the DOM would be in an indeterminate state. To give the script a consistent, complete view of the document, the parser can only stop and hand control to the script.
The cost is direct:
- Downloading the script takes time (longer on a slow network);
- Executing the script takes time (longer if there's a heavy loop);
- During all of it, HTML parsing is completely halted, later content doesn't appear, and the page is blank.
This answers the question from the start: why an infinite-loop <script> can stop the whole page from appearing — because the parser is parked right there waiting for it to finish, and it never will.
5. Three Script-Loading Strategies: Plain / async / defer
Since plain scripts block parsing, the browser gives us two switches to change that behavior:
<script src="a.js"></script> <!-- plain: blocks parsing -->
<script async src="b.js"></script> <!-- async: download doesn't block; executes as soon as ready -->
<script defer src="c.js"></script> <!-- defer: download doesn't block; executes in order after parsing -->Compared:
| Form | Does download block parsing? | When it runs | Order guaranteed? |
|---|---|---|---|
<script> | Blocks (parsing halts during download + execution) | Immediately | Yes (because it blocks) |
<script async> | No | As soon as it downloads (may interrupt parsing) | No |
<script defer> | No | After HTML parsing, before DOMContentLoaded | Yes, source order |
A few practical guidelines:
asyncsuits independent scripts (analytics, ads). It runs as soon as it's downloaded, may interrupt parsing, and its order is unpredictable.deferis more predictable: it doesn't block parsing and guarantees source order plus execution after the DOM is ready. Most modern scripts are fine withdefer.- Module scripts
<script type="module">default todeferbehavior, which is one more reason modern projects favor modules.
Remember the difference in one line: async is "run the moment I'm downloaded, ignore everyone else"; defer is "queue up, and run in order after HTML parsing."
6. Where CSS Fits In: The CSSOM and "Render Blocking"
Besides <script>, the parser also runs into styles:
<style>...</style>and<link rel="stylesheet">are parsed and built into another tree — the CSSOM (CSS Object Model).
Here's the commonly confused part:
- External CSS does not block HTML parsing (the parser keeps building the DOM).
- But CSS blocks rendering. The browser doesn't want you to see a "naked, unstyled page" that then suddenly becomes pretty (that flash is called FOUC). So it waits for the CSSOM before the first render.
- Subtler still: CSS blocks the execution of subsequent
<script>s. Because a script might need to read styles (e.g.,getComputedStyle), the browser must ensure the CSSOM is ready before letting the script run. So CSS → holds up the script → indirectly holds up parsing.
So "CSS doesn't block" is wrong. Precisely: CSS doesn't block DOM construction, but it blocks rendering and blocks subsequent scripts.
7. Completion Signals: DOMContentLoaded vs load
When this chain finishes, it fires two often-confused events:
DOMContentLoaded: fires when the DOM tree has finished parsing (anddeferscripts have executed). Images, styles, and iframes may not have loaded yet.load: fires when all resources (images, styles, fonts, iframes…) have finished loading.
This explains a classic pattern:
document.addEventListener("DOMContentLoaded", () => {
// the DOM is guaranteed ready here; safe to query/manipulate nodes
});If you put a script in <head> without defer, it runs before body has been parsed, so querying nodes yields null — which is exactly where the advice "put scripts at the bottom, or add defer, or wrap in DOMContentLoaded" comes from.
8. Back to the Developer's View
Lining the chain up, a lot of behaviors suddenly make sense:
- "Why put scripts before
</body>?" So parsing can get through the earlier content first, letting the first screen render before scripts run. - "Why use
defer/async?" To take "downloading the script" off the parsing path, so it doesn't block the HTML. - "Why avoid big JS on the first screen?" Plain script execution halts parsing; big JS = a long blank screen.
- DevTools: the Performance panel shows
Parse HTMLandEvaluate Scriptblocks, and how they hold up later rendering.
Deeper Understanding (Common Misconceptions)
- "The DOM tree is just my HTML source." No. The parser completes, moves, and even discards nodes (
<p>auto-closing, table foster parenting). The Elements panel shows the parsed DOM. - "HTML only starts parsing once it's fully downloaded." No. It's incremental/streaming: build as you download.
- "
asyncis always better thandefer." Not necessarily.asyncdoesn't guarantee order and may interrupt parsing and rendering;deferis more predictable. Which one to use depends on whether the script is independent. - "CSS doesn't block." It doesn't block DOM construction, but it blocks rendering and blocks subsequent scripts.
- A fun fact: the preload scanner. While parsing is stuck on a script, modern browsers aren't actually idle — they have a separate "preload scanner" that quickly looks ahead, discovers resources like
<img>,<link>, and later<script>s, and starts downloading them early. So a script blocks "tree construction," but not necessarily "resource downloads." This is why a browser can be blocked by a script while images have already been fetched.
Hands-On Practice
Basic Exercises
-
See the DOM differ from the source
Prepare a file:
html<p>first<p>second <table><div>misfiled</div></table>Open it in a browser and look at the structure the browser builds in the Elements panel —
</p>is added for you, and thedivis moved out of thetable. -
Manufacture a parse block yourself
Write a page with an inline script in the middle of the content:
html<h1>this line appears</h1> <script> const t = Date.now(); while (Date.now() - t < 3000) {} // stuck for 3 seconds </script> <h1>this line waits 3 seconds to appear</h1>You'll see: the first line appears quickly (it was already parsed and rendered), and the second waits until the script finishes — a direct feel for parsing being paused by a script.
-
Compare
asyncanddeferorderWrite three scripts that each
console.logthemselves, one plain, oneasync, onedefer. Refresh a few times and watch the output order.
Advanced Challenge
- Open DevTools → Performance and record a page load. Find the
Parse HTMLand script-execution time blocks, and see how a plain script in<head>pushes the first render later. Then change it todeferand compare.
Summary & Keywords
- HTML's journey is: bytes → decode → characters → tokenize → tokens → tree construction → DOM tree, and it's an incremental pipeline.
- The HTML parser is error-tolerant: it completes and moves nodes, so the DOM tree is not necessarily equal to the source.
- A plain
<script>pauses HTML parsing, because it may change the document or may depend on "the complete DOM up to that point." async= run immediately once downloaded, no order guarantee;defer= run in order after parsing; module scripts default todefer.- CSS doesn't block DOM construction, but blocks rendering and blocks subsequent scripts.
DOMContentLoaded= DOM ready;load= all resources ready.
Keywords: HTML parser, tokenization, tree construction, DOM, incremental/streaming parsing, error tolerance, parser-blocking, async, defer, CSSOM, DOMContentLoaded, load, preload scanner
Further Reading and Next Article Preview
At this point the browser holds two key things: the DOM tree (content) and the CSSOM (styles). But both are still just "data structures" — nothing is on screen yet.
Next, we continue downward: how does the browser turn DOM + CSSOM into the pixels you see? That's the rendering pipeline — style computation (Style) → layout (Layout) → paint (Paint) → composite (Composite) — and why "changing a width" and "changing a transform" cost wildly different amounts.
Earlier articles:
