The HTML Parser: From Bytes to the DOM Tree
In the previous article we watched the main thread and its scheduler from a distance: we know rendering is the third step of the event loop, and that synchronous JavaScript blocks rendering.
But a more fundamental question has been hanging in the air: when the browser receives a chunk of HTML, what actually happens?
"View page source" shows you a bunch of <div>s, <p>s and text. Yet in JavaScript it becomes a DOM tree you can query with document.querySelector('.title'). Who performs that translation?
Even more counterintuitive: drop a <script> containing an infinite loop into a page, and the whole thing goes blank and freezes; and a slow stylesheet in <head> can make the content take ages to show up. How can a bit of JavaScript or CSS dictate when the whole page appears?
In this article we'll follow the browser from the bytes on the network all the way to the DOM tree you manipulate — and along the way, settle two questions that are usually hand-waved: "why does CSS block rendering / block scripts", and "what is the first paint actually waiting for?" Once you see this chain, a lot of "truisms" (put scripts at the bottom, use defer, inline critical CSS) stop being things to memorize.
What You'll Learn
- How a chunk of HTML goes from "bytes" to "DOM tree" step by step
- Why the HTML parser is "incremental" (parses while downloading) and "error-tolerant" (fixes broken markup)
- Why a plain
<script>pauses HTML parsing - What actually differs between
asyncanddefer - How the browser knows a CSS file is "done", and how the CSSOM grows piece by piece
- What "CSS blocks rendering" and "CSS blocks scripts" each mean (the most commonly confused pair)
- Why the first paint never waits for a "complete" DOM or CSSOM
- The difference between
DOMContentLoadedandload
Prerequisites
- It helps to read the event loop article first, to understand the main thread and "rendering is step 3 of the loop":
- No compiler-theory background needed; this starts from zero
Starting from a Scenario
Suppose the server sends the browser this HTML:
<!DOCTYPE html>
<html>
<head>
<link rel="stylesheet" href="style.css">
</head>
<body>
<h1 class="title">Hello</h1>
<p>world</p>
<script src="app.js"></script>
<p>after script</p>
</body>
</html>The browser ultimately needs to end up with a tree it can manipulate:
Document
└── html
├── head
│ └── link
└── body
├── h1.title ("Hello")
├── p ("world")
├── script
└── p ("after script")Now the questions:
- What the network gave the browser was just a sequence of bytes — how did it become this tree?
- If
app.jscontains awhile(true){}, why does even<p>after script</p>fail to appear?
The answer to the second question will explain a whole family of performance behaviors we take for granted.
Core Content
1. One Pipeline: Bytes → Characters → Tokens → DOM Tree
What the browser receives as HTML is fundamentally a byte stream (a pile of 0s and 1s). It isn't text, and it certainly isn't a tree. To become the DOM, it goes through a few steps:
network → byte stream
│ ① decode (using the declared encoding, usually UTF-8)
▼
character stream <html><body>hello</body>...
│ ② tokenize
▼
token sequence {StartTag: html} {StartTag: body} {Text: "hello"} {EndTag: body}
│ ③ tree construction
▼
DOM tree (a tree of nodes)- Decode: bytes have no "character meaning" by themselves. The browser determines the encoding from hints like the BOM,
Content-Type, and<meta charset>, and turns bytes back into characters (this step is called "encoding sniffing" in the HTML standard). - Tokenize: cut the character stream into meaningful tokens — start tags, end tags, text, comments, the DOCTYPE, and so on.
- Tree construction: assemble those tokens into a tree following their nesting — a start tag hangs a node below, an end tag moves back up.
The key point: these three steps are not a three-stage process where everything is decoded, then everything is tokenized, then everything is built. They're one pipeline — a little bytes come in, and it advances a little. Understanding this is the key to the next section on "incremental" parsing.
2. Incremental Parsing: Parse While You Download
The network doesn't hand the whole HTML to the browser at once; it arrives in chunks. So the browser does something bold and clever: the parser doesn't wait for the HTML to arrive complete — it parses as it comes.
chunk 1 received: <div><p>hello
→ the parser can already create the div and p nodes and put in some text
chunk 2 received: </p></div><span>...
→ keep on buildingThis is incremental parsing, also called streaming parsing. It has a few direct consequences:
- Faster first paint: the browser doesn't need the full 1 MB of HTML before it can start rendering the beginning.
- You can touch the DOM mid-parse: that's why you can attach events to
document.bodyand script can modify nodes that have already appeared.
Think about the alternative: if the browser insisted on the full download before building the tree, a large page's first paint would wait a long time. Incremental parsing is one of the foundations of the web's feel.
3. Error-Tolerant Parsing: HTML Is Not "Strict Mode"
If you've written XML, you know how strict it is: one unclosed tag or misplaced nesting and it errors out and refuses to work.
HTML is the opposite — its parsing rules are error-tolerant. Faced with broken, scrambled markup, the browser doesn't error; it "repairs" it according to an established set of rules and renders the page as best it can. A few examples you'll run into:
you wrote: <p>first<p>second
browser reads: <p>first</p><p>second</p> ← closes tags for you
you wrote: <table><div>misfiled</div></table>
browser reads: move the div outside the table (foster parenting) ← builds structure for youWhy is HTML so forgiving? Because the web must be backward compatible. Historically, so many pages were written incorrectly that if browsers crashed on the first mistake like XML does, those pages would have become unusable. So the standards authors chose to "render as best as possible."
This leads to an important fact: the source you write and the DOM tree you get are not necessarily the same shape. The parser may complete, move, or even discard some of your nodes. What DevTools' Elements panel shows is the parsed DOM, not your source.
4. The Crucial Part: Why <script> Pauses Parsing
Now we reach the most important section.
When the parser encounters a plain <script> (with neither async nor defer) in the HTML, it does something that looks brutal:
pause HTML parsing
→ download and immediately execute this JavaScript
→ when it finishes, resume parsingThis is what parser-blocking means. Why stop? Not because the browser is lazy, but because a script might change the document in turn:
- The script might call
document.write(...)to insert HTML into the document; - The script might query the DOM (
document.querySelector(...)) and expects to see "the fully parsed DOM up to this point."
If the browser kept parsing while the script freely changed things, the DOM would be in an indeterminate state. To give the script a consistent, complete view of the document, the parser can only stop and hand control to the script.
The cost is direct:
- Downloading the script takes time (longer on a slow network);
- Executing the script takes time (longer if there's a heavy loop);
- During all of it, HTML parsing is completely halted, later content doesn't appear, and the page is blank.
This answers the question from the start: why an infinite-loop <script> can stop the whole page from appearing — because the parser is parked right there waiting for it to finish, and it never will.
One easily-glossed-over point: if a stylesheet before this script hasn't finished loading, the "pause" gets even longer. But remember — the pause is caused by the script; CSS only makes that pause longer. Section 7 clears this up specifically.
5. Three Script-Loading Strategies: Plain / async / defer
Since plain scripts block parsing, the browser gives us two switches to change that behavior:
<script src="a.js"></script> <!-- plain: blocks parsing -->
<script async src="b.js"></script> <!-- async: download doesn't block; executes as soon as ready -->
<script defer src="c.js"></script> <!-- defer: download doesn't block; executes in order after parsing -->Compared:
| Form | Does download block parsing? | When it runs | Order guaranteed? |
|---|---|---|---|
<script> | Blocks (parsing halts during download + execution) | Immediately | Yes (because it blocks) |
<script async> | No | As soon as it downloads (may interrupt parsing) | No |
<script defer> | No | After HTML parsing, before DOMContentLoaded | Yes, source order |
A few practical guidelines:
asyncsuits independent scripts (analytics, ads). It runs as soon as it's downloaded, may interrupt parsing, and its order is unpredictable.deferis more predictable: it doesn't block parsing and guarantees source order plus execution after the DOM is ready. Most modern scripts are fine withdefer.- Module scripts
<script type="module">default todeferbehavior, which is one more reason modern projects favor modules.
Remember the difference in one line: async is "run the moment I'm downloaded, ignore everyone else"; defer is "queue up, and run in order after HTML parsing."
6. How CSS Gets Into the Tree: The CSSOM Is "Live", and How the Browser Knows It's Ready
So far we've talked about the DOM. Now let's bring CSS in.
Rendering uses two trees: the DOM (content) and the CSSOM (styles — the CSS Object Model). We know where the DOM comes from; where does the CSSOM come from, and when does it count as "ready"?
6.1 HTML parsing and CSS parsing don't block each other
While the parser keeps building the DOM, styles are downloading and being parsed on a separate track. External CSS does not stop HTML parsing — something many people assume happens but doesn't.
6.2 The CSSOM is an incrementally-built "live" structure
The CSSOM isn't built all at once once every style has arrived; it adds rules to one table as it parses along. Multiple stylesheets merge into the same CSSOM, and the order in which they're added determines the cascade priority.
Styles come from three sources, with different ready times:
inline <style> → when the parser reaches it, parsed into the CSSOM on the spot (synchronous, immediately ready)
external <link> → starts downloading on encounter; rules are added only after download + parse (asynchronous)
JS-inserted:
├─ insertRule(...) → modifies the CSSOM directly (synchronous, immediately effective)
└─ createElement('link') → loads asynchronously; effective once loaded6.3 How the browser decides "this CSS is done"
There is no global "CSS is OK" signal. What the browser maintains is a "set of stylesheets not yet ready": when a stylesheet becomes ready, it's removed from the set. And "ready" is decided by resource-level completion signals — the network receiving the full resource plus the stylesheet being parsed into the CSSOM (an external <link> also fires its own load event). It has nothing to do with "how far DOM parsing has gotten."
6.4 The CSSOM is never "final"
Because JS can modify it at any time (insertRule synchronously, inserting a <link> asynchronously), the CSSOM is a live table. This leads to an important conclusion: the browser can only wait for styles it already knows about. A stylesheet inserted in the future is unknown to it right now, so there's nothing to wait for — once it's actually inserted and the CSSOM changes, that triggers a new render.
7. Two Kinds of "Blocking" — Don't Mix Them Up
CSS gives rise to two kinds of "blocking" that sound similar but are entirely different.
7.1 CSS blocks rendering (render-blocking)
Until (known, and render-affecting) CSS is ready, the browser won't paint the page.
Reason: the first step of rendering is style computation — DOM + CSSOM → render tree. Without a ready CSSOM, that step can't happen, so it can't paint. The browser would rather not paint at all than show you a naked, unstyled page that then suddenly changes shape (that flash is called FOUC).
The key point: CSS only blocks "rendering / first paint," not "DOM construction." HTML parsing continues, the DOM keeps growing; the screen just stays blank.
7.2 CSS blocks scripts (stylesheets block scripts)
If a
<script>has a not-yet-loaded stylesheet before it, the browser waits for that stylesheet before executing the script.
Reason: the script might read style-related information (getComputedStyle(el).color, el.offsetHeight…), whose values depend on CSS. To avoid the script reading styles that haven't taken effect yet, the browser must have the CSSOM ready before letting it run.
7.3 The key: don't invert the cause
This is the easiest thing to get wrong: HTML parsing stops because of the <script>, not because of CSS. Keep the two separate:
"script blocks parsing" — the moment the parser meets a plain <script>, parsing stops (nothing to do with CSS)
"CSS blocks scripts" — a script's "prerequisites to run" gains an extra item: "wait for the stylesheets before me"In other words: the parser meets <script> → parsing stops (the script did this); and before the script actually runs, it may have to wait extra for preceding stylesheets to be ready (this part is CSS's doing). CSS "lengthens the pause" — it doesn't "cause the pause."
Two contrasting examples make it clear.
External CSS up front, still loading:
<link rel="stylesheet" href="slow.css"> <!-- not downloaded yet -->
<script> console.log("hi"); </script> <!-- must wait for slow.css's CSSOM -->→ Parsing stops at <script>; the script first waits for slow.css; then it runs; only then does parsing resume. The pause was lengthened by CSS.
Inline CSS + inline script:
<style> h1 { color: red; } </style> <!-- ready synchronously -->
<script> console.log("hi"); </script> <!-- no "unready" style ahead, so nothing to wait for -->→ The inline <style> enters the CSSOM the instant the parser reaches it, so the script doesn't wait for CSS at all; it simply pauses parsing "because it's a plain script" and then runs immediately. In this example there is no CSS wait whatsoever.
8. What the First Paint Actually Waits For: Not "Complete", but "Ready + Something to Paint"
Section 7 said "CSS blocks rendering," which is easily read as "the browser waits for all CSS and a fully-parsed DOM before rendering." That's not it. The key sentence is:
Rendering is not a one-time event; it happens over and over. From page open to close, the browser renders hundreds or thousands of times. Each time, it uses "the DOM and CSSOM at that instant", not "the final ones."
So for "one frame," there are only two conditions:
① stylesheets that have been discovered and that affect this render → all ready (not "all CSS in the world")
② enough content on hand to paint anything
Not required:
✗ the entire DOM parsed
✗ all CSS (including stylesheets yet to appear) arrived- Why the DOM needn't be complete: parsing is incremental, and rendering is progressive too. The top of the page can appear first, with the rest growing in — direct evidence that "the DOM isn't complete yet the page has already rendered."
- Why the CSSOM needn't be complete: the browser can only wait for styles it already knows about. A
<link>encountered later in the document, or a style JS inserts later, is unknown at first paint, so it isn't waited for; when it arrives, it triggers the next frame.
A handy analogy: rendering is like taking snapshots continuously, not waiting for the whole world to be in place before taking one. But before snapping, it confirms that "the lights it already knows about are on," so it doesn't take an ugly, unstyled (FOUC) photo.
This also lets you answer a question you might have: "If I insert a CSS file 100 seconds later via setTimeout, will the browser wait for it?" — No. Before insertion it has no idea the CSS exists; once it's inserted and the CSSOM changes, that triggers one more rendered frame.
9. Completion Signals: DOMContentLoaded vs load
When this chain finishes, it fires two often-confused events:
DOMContentLoaded: fires when the DOM tree has finished parsing (anddeferscripts have executed). Images, styles, and iframes may not have loaded yet.load: fires when all resources (images, styles, fonts, iframes…) have finished loading.
This explains a classic pattern:
document.addEventListener("DOMContentLoaded", () => {
// the DOM is guaranteed ready here; safe to query/manipulate nodes
});If you put a script in <head> without defer, it runs before body has been parsed, so querying nodes yields null — which is exactly where the advice "put scripts at the bottom, or add defer, or wrap in DOMContentLoaded" comes from.
10. Back to the Developer's View
Lining the chain up, a lot of behaviors suddenly make sense:
- "Why put scripts before
</body>?" So parsing can get through the earlier content first, letting the first screen render before scripts run. - "Why use
defer/async?" To take "downloading the script" off the parsing path, so it doesn't block the HTML. - "Why avoid big JS on the first screen?" Plain script execution halts parsing; big JS = a long blank screen.
- "Why inline critical CSS?" Because CSS blocks rendering: the later the first-screen styles arrive, the longer the blank screen; inlining them into the HTML saves a round trip.
- DevTools: the Performance panel shows
Parse HTML,Evaluate Script, andRecalculate Styleblocks, and how they hold up later rendering.
Deeper Understanding (Common Misconceptions)
- "The DOM tree is just my HTML source." No. The parser completes, moves, and even discards nodes (
<p>auto-closing, table foster parenting). The Elements panel shows the parsed DOM. - "HTML only starts parsing once it's fully downloaded." No. It's incremental/streaming: build as you download.
- "
asyncis always better thandefer." Not necessarily.asyncdoesn't guarantee order and may interrupt parsing and rendering;deferis more predictable. Which one to use depends on whether the script is independent. - "CSS doesn't block." It doesn't block DOM construction, but it blocks rendering and blocks subsequent scripts.
- "Rendering waits for a complete DOM/CSSOM." No. Rendering only uses the two trees "as they exist right now"; it waits for "known, not-yet-ready, render-affecting stylesheets" plus "enough content to paint." Rendering happens over and over.
- "CSS inserted in the future will be waited for." It won't. The browser can't predict the future; only resources it has already discovered take part in "waiting." Late CSS merely triggers a re-render.
- A fun fact: the preload scanner. While parsing is stuck on a script, modern browsers aren't actually idle — they have a separate "preload scanner" that quickly looks ahead, discovers resources like
<img>,<link>, and later<script>s, and starts downloading them early. So a script blocks "tree construction," but not necessarily "resource downloads." This is why a browser can be blocked by a script while images have already been fetched.
Hands-On Practice
Basic Exercises
-
See the DOM differ from the source
Prepare a file:
html<p>first<p>second <table><div>misfiled</div></table>Open it in a browser and look at the structure the browser builds in the Elements panel —
</p>is added for you, and thedivis moved out of thetable. -
Manufacture a parse block yourself
Write a page with an inline script in the middle of the content:
html<h1>this line appears</h1> <script> const t = Date.now(); while (Date.now() - t < 3000) {} // stuck for 3 seconds </script> <h1>this line waits 3 seconds to appear</h1>You'll see: the first line appears quickly (it was already parsed and rendered), and the second waits until the script finishes — a direct feel for parsing being paused by a script.
-
Compare
asyncanddeferorderWrite three scripts that each
console.logthemselves, one plain, oneasync, onedefer. Refresh a few times and watch the output order.
Observing CSS and Rendering
-
See the CSSOM exist for yourself
Open any page's console and type:
jsdocument.styleSheets;You'll get a (live) list — that's the CSSOM the browser has already built. Poke at
document.styleSheets[0].cssRulesto see the actual rules. -
Watch a late-arriving CSS trigger a re-render
Open a blank white page and run this in the console:
jssetTimeout(() => { document.head.insertAdjacentHTML( "beforeend", "<style>body { background: #222; color: #fff; }</style>", ); }, 3000);For the first 3 seconds the page stays as-is (the browser does not wait for this not-yet-existing CSS), then the style is inserted and the page "flashes" to a dark theme — proof that "late CSS merely triggers the next frame."
Advanced Challenge
- Open DevTools → Performance and record a page load. Find the
Parse HTML,Evaluate Script, andRecalculate Styletime blocks, and see how a plain script in<head>— and a slow stylesheet — each push the first render later. Then switch them todeferand inline CSS respectively, and compare.
Summary & Keywords
- HTML's journey is: bytes → decode → characters → tokenize → tokens → tree construction → DOM tree, and it's an incremental pipeline.
- The HTML parser is error-tolerant: it completes and moves nodes, so the DOM tree is not necessarily equal to the source.
- A plain
<script>pauses HTML parsing, because it may change the document or may depend on "the complete DOM up to that point." async= run immediately once downloaded, no order guarantee;defer= run in order after parsing; module scripts default todefer.- The CSSOM is incremental and live; the browser tracks a "set of not-yet-ready stylesheets" via resource-completion signals to decide whether CSS is ready, not via "is the DOM parsed yet."
- Keep the two kinds of blocking apart: CSS blocks rendering (paint only once DOM+CSSOM are ready, to prevent FOUC), and CSS blocks scripts (a script waits for the stylesheets before it). And parsing stops because of the
<script>; CSS merely lengthens that pause. - The first paint never waits for "complete": it uses "the DOM and CSSOM right now," waiting only for "known, not-yet-ready styles + something to paint," so rendering is progressive; late DOM/CSS merely triggers a new render.
DOMContentLoaded= DOM ready;load= all resources ready.
Keywords: HTML parser, tokenization, tree construction, DOM, CSSOM, incremental/streaming parsing, error tolerance, parser-blocking, render-blocking, stylesheets block scripts, progressive rendering, async, defer, FOUC, DOMContentLoaded, load, preload scanner
Further Reading and Next Article Preview
Before we get to "pixels," there's a more fundamental question that deserves its own article: the HTML parser we just covered, and the event loop from earlier — are they two things, or one thing described in two ways? The parser is a "straight line," the event loop is a "circle" — how do they fit together?
The next article welds the two together — the parser isn't a separate loop; it's a series of tasks on the event loop:
Only after that do we take the final step down: how does the browser turn DOM + CSSOM into the pixels you see? That's the rendering pipeline — style computation (Style) → layout (Layout) → paint (Paint) → composite (Composite) — and why "changing a width" and "changing a transform" cost wildly different amounts.
Earlier articles:
