From bytes to pixels.
A web page reaches the browser as a stream of bytes and leaves it as a grid of coloured pixels, redrawn up to sixty or more times a second while something on it moves. In between sits an assembly line: parse, style, lay out, paint, rasterize, composite. Each station hands the next a different data structure, and each one can be skipped when nothing it depends on has changed. Once you can name the stations and say which of them a change wakes up, most front-end performance advice stops being a list of rules and starts being common sense.
The idea.
Rendering is not one step. It is a chain of conversions, each producing the input of the next, and the browser reruns only the links a change actually touches.
Open Copperpot's product page for the Enamel casserole on a phone. The server sends text: 96 KB of markup and, a little later, a 48 KB stylesheet. The screen needs something completely different, a rectangle of pixels with a colour for each one. No single algorithm goes straight from one to the other. The browser gets there in stages, and each stage builds a structure that answers one question.
Parsing asks what the document contains, and answers with the DOM. Parsing the stylesheet asks which rules exist, and answers with the CSSOM. Style asks which of those rules win for each element, and answers with a computed style per element. The layout tree asks which elements actually produce boxes. Layout asks where each box goes and how big it is. Paint asks what to draw and in which order, and answers with a list of drawing commands. Raster runs those commands to get pixels, and compositing stacks the rasterized layers into the frame you see.
The split is what makes the web fast enough. Because every stage keeps its output, a later change only has to redo the stages whose inputs it changed. Recolouring a button border does not move anything, so layout is skipped. Sliding a photo with a transform does not even need repainting. Knowing which stage a change wakes up is the single most useful thing this topic teaches.
The pipeline, first load
What each stage produces for Copperpot's page
| Stage | Input | Output | On Copperpot's product page |
|---|---|---|---|
| Parse HTML | Bytes of markup | DOM | About 1,200 element nodes plus their text nodes |
| Parse CSS | Bytes of stylesheet | CSSOM | About 1,900 rules from main.css, plus the browser's own defaults |
| Style | DOM + CSSOM | Computed style per element | Every element gets a value for every property, inherited or not |
| Layout tree | DOM + computed styles | Tree of boxes | Hidden size guide dropped; ::before badges added |
| Layout | Layout tree + viewport width | Sizes and positions | Recomputed if the phone is rotated: 393 px wide becomes 852 px |
| Paint | Geometry + styles | Display lists | Commands such as fill rectangle, draw text run, draw image |
| Raster | Display lists | Pixels in tiles | Image decode of the hero photo happens here |
| Composite | Rasterized layers | A frame | Header, page and toast stacked in the right order |
How it works.
The same Copperpot page, one station at a time, with what each one reads, what it builds and what it refuses to do until its inputs are ready.
Bytes to characters to tokens
Before anything can be parsed, bytes have to become characters, and that needs an encoding. The browser looks for a byte order mark at the very start, then for a charset on the Content-Type header, and failing both it scans the beginning of the file for a meta charset declaration. The HTML Standard encourages browsers to look only at the first 1,024 bytes for it, so Copperpot puts <meta charset="utf-8"> as the first element in head. Declared later, or not at all, a page can be decoded with a guessed encoding and, when the guess turns out wrong, decoded again from the start.
The tokenizer then reads characters and emits tokens: a doctype, start tags with their attributes, end tags, runs of text, comments. It is a state machine that changes mode as it goes, so the same less-than sign means a tag in body text, plain text inside a <textarea> and nothing special inside a <script> until the matching end tag. Tokens are handed to tree construction one at a time; the browser does not wait for the whole file.
The top of Copperpot's product page, and the tree it becomes
<!doctype html>
<meta charset="utf-8">
<title>Enamel casserole, 24 cm · Copperpot</title>
<link rel="stylesheet" href="/css/main.css">
<header class="site-bar">
<a href="/" class="logo">Copperpot</a>
<button class="basket" aria-label="Basket, 0 items"></button>
</header>
<main class="product">
<h1>Enamel casserole, 24 cm</h1>
<p class="price">£64
<p class="stock">In stock, ships Monday
<table class="specs">
<tr><td>Capacity<td>4.2 litres
#document
├─ <!DOCTYPE html>
└─ html ← never written, inserted
├─ head ← never written, inserted
│ ├─ meta charset=utf-8
│ ├─ title "Enamel casserole, 24 cm · Copperpot"
│ └─ link rel=stylesheet
└─ body ← inserted at the first <header>
├─ header.site-bar …
└─ main.product
├─ h1 "Enamel casserole, 24 cm"
├─ p.price "£64" ← closed when the next <p> opened
├─ p.stock "In stock, ships Monday"
└─ table.specs
└─ tbody ← never written, inserted
└─ tr
├─ td "Capacity"
└─ td "4.2 litres"
Tokens to the DOM
Tree construction keeps a stack of open elements and decides, token by token, where each node goes. It knows the rules a person writing the markup may have skipped: html, head and body exist even when nobody wrote them, a table row always sits inside a tbody, and a new <p> closes an open one. Copperpot's template omits all of those and the tree still comes out the same in every browser, because the standard defines a recovery for every kind of mistake. There is no such thing as a fatal error when parsing HTML, which is very different from parsing JSON or XML.
Nodes are appended as soon as their tokens arrive, so the DOM grows while the bytes are still downloading. That is what lets a browser show the header and the title of a long page before its footer exists. Two things can interrupt the flow: a classic script without async or defer makes the parser stop and run it, because the script could write more markup into the stream, and the first paint waits for the stylesheets in head. Who waits for whom, and how async, defer and the preload scanner change it, is the subject of What blocks the parser and the first paint.
Markup Copperpot's template gets wrong, and what the parser does
| Written | Parser's recovery | Visible effect |
|---|---|---|
| <p class="price">£64 then <p class="stock"> | The second start tag closes the open paragraph first | Two sibling paragraphs, as intended |
| <table><tr> | Inserts a tbody element around the row | A selector table > tr matches nothing; table > tbody > tr does |
| <p>See <div>size guide</div></p> | A div cannot sit inside a p, so the p is closed before it and the stray </p> becomes an empty paragraph | An extra empty p after the div, with its margins |
| <b>Sale <i>today</b> only</i> | Closes and reopens the formatting elements so the tree stays a tree | "only" stays italic but not bold |
| No <html>, <head> or <body> | Inserts all three at the right moments | None; document.body still works |
The CSSOM and style
main.css goes through its own parser and becomes the CSSOM: style sheets holding rules, each rule a selector and a block of declarations. It is a separate structure from the DOM, and on its own it says nothing about any particular element.
Style calculation joins the two. For each element the engine finds the rules whose selectors match it, sorts the competing declarations by the cascade (origin and importance, then layers, then specificity, then order of appearance), fills in what nothing set by inheriting from the parent or using the initial value, and resolves relative values where it can, such as 1.5em to 24px. The result is a computed style for every element and every property. Copperpot's price paragraph ends up with a colour, a font size and a margin even though its own rule only names two of those.
Engines avoid repeating work here: they index rules so that most selectors are never tried against most elements, share computed styles between siblings that match the same rules, and after a change restyle only the elements a changed class or attribute could affect. Style is still proportional to how many elements need it, which is why a class toggled on body can be more expensive than the same class on one button.
Where p.price's computed values come from
| Property | Computed value | Came from |
|---|---|---|
| font-size | 24px | .product .price { font-size: 1.5rem } with a 16px root |
| color | rgb(122, 44, 20) | .price { color: var(--rust) }, the custom property resolved |
| font-family | "Fraunces", serif | Inherited from main.product, which inherited it from body |
| margin-top | 24px | p { margin-block: 1em } in the browser's own stylesheet; em uses the element's own font size, so 1em is 24px here |
| display | block | The browser's default for p |
From elements to boxes
Layout does not work on the DOM. It works on a tree of boxes built from the DOM and the computed styles, often called the layout tree or render tree. The two trees disagree in instructive places. Elements whose display is none produce no box at all, and neither do their descendants; head and everything in it is display:none by default. Pseudo-elements such as ::before produce boxes even though they are not in the DOM. An element with display:contents gives up its own box but its children keep theirs, as if they belonged to its parent. And an element with visibility:hidden keeps its box, because it still takes up space; it is simply not drawn.
Part of Copperpot's product block, as nodes and as boxes
Four ways to hide Copperpot's size guide, and what each leaves in the pipeline
| CSS | In the layout tree? | Takes space? | Painted? | Screen readers |
|---|---|---|---|---|
| display: none | No | No | No | Removed, with all its descendants |
| visibility: hidden | Yes | Yes | No | Not announced |
| opacity: 0 | Yes | Yes | Yes, fully transparent | Still announced, and still clickable |
| content-visibility: hidden | Its own box only; the contents are skipped | Its own size (set it with contain-intrinsic-size) | Contents no | Contents hidden; their rendering state is kept for a quick reveal |
Layout
Layout walks the box tree and works out the size and position of every box. Widths mostly flow down from the parent: the viewport is 393 CSS pixels wide on Copperpot's test phone, main has 16px of padding on each side, so the product block's content is 361px wide. Heights mostly flow up from the children: the price paragraph is as tall as its line of text, which depends on the font, which may not have loaded yet. Text is shaped and broken into lines here, so a longer product name can push everything below it down. The rules that decide where each box goes (normal flow, flex, grid, positioned boxes) belong to the formatting contexts topics; what matters here is that layout's output is geometry, and that geometry is what paint and hit testing read.
Because a box's position depends on what comes before it and its size on what is inside it, a change in one place can move many others. Engines mark dirty boxes and skip clean subtrees, and in Chromium the result of layout is an immutable fragment tree that is reused where its inputs did not change. The first layout of a page is usually just called layout; later ones are often called reflows.
Paint, then raster
Paint does not produce pixels. It walks the laid-out boxes in paint order and records commands: fill this rounded rectangle with cream, draw this text run in rust at these coordinates, draw this image into that rectangle. Paint order is not DOM order. Backgrounds go under borders, borders under content, and positioned or z-indexed boxes go where their stacking context puts them, so Copperpot's sticky header is recorded after the product photo it overlaps even though the header comes first in the markup.
Raster executes those commands into actual pixels. Large areas are cut into tiles and drawn in parallel on raster threads, often with the GPU's help, starting with the tiles in or near the viewport. Images are decoded here too, which is why a large hero photo can sit in memory as compressed bytes well before it is ready to draw. Keeping raster off the main thread means a slow decode does not freeze script or input.
Composite
Parts of a page may be painted into separate layers: a video, a canvas, a fixed header, an element being animated with transform or opacity. Compositing stacks those rasterized layers with their current offsets, transforms, clips and opacities, and the result is the frame sent to the display. Because the layers are already pixels, moving one or fading it means recompositing, not repainting. That is how a page can scroll, or a photo can slide, while the main thread is busy elsewhere. Which elements get a layer of their own, and what that costs in memory, is covered in Layers and compositing.
The clock the pipeline races
The whole first load, on one clock
Copperpot's product page from the first byte to the hero photo
Scenario 1 of 2: As described.
Timeline as a list
Copperpot's product page from the first byte to the hero photo: 6 lanes, from 0 ms to 360 ms.
- 0–8 ms · Network · HTML chunk 1
- 2–28 ms · Main: parse HTML · head, header, title, price
- 4–180 ms · Network · main.css in flight
- 4 ms · Main: parse HTML · finds main.css
- 8–300 ms · Network · hero photo in flight
- 28–180 ms · Screen, Main: style, layout, paint · window: DOM ahead of the CSSOM
- 58–66 ms · Network · chunk 2
- 66–84 ms · Main: parse HTML · gallery, options
- 108–114 ms · Network · chunk 3
- 114–126 ms · Main: parse HTML · reviews, footer
- 126 ms · Main: parse HTML · DOM done: interactive (ok)
- 180–186 ms · Main: style, layout, paint · CSSOM
- 180 ms · Network · main.css arrives
- 186–200 ms · Main: style, layout, paint · style
- 200–214 ms · Main: style, layout, paint · layout
- 214–221 ms · Main: style, layout, paint · paint
- 221–224 ms · Compositor thread · commit
- 224–232 ms · Raster and decode · raster tiles
- 233.3 ms · Screen · FP and FCP: first frame (ok, ok)
- 300–303 ms · Main: style, layout, paint · paint image
- 303–318 ms · Raster and decode · decode + raster hero
- 333.3 ms · Screen · hero shown (ok)
Two things to read off the waterfall. First, parsing is not the bottleneck on this page: the DOM is complete 54 ms before the stylesheet arrives, and the screen stays blank for that whole window. Second, rendering is incremental when its inputs allow it. With the CSSOM ready early, the browser paints the part of the DOM it has and adds to it as chunks arrive, so the visitor sees the price a sixth of a second sooner. Two milestones from the Paint Timing specification name these moments: first paint is the first frame that shows anything other than the default background, and first contentful paint is the first frame that shows text, an image, a non-blank canvas or an SVG from the DOM. On this page the first frame already holds the product name, so the two coincide; a page whose first frame is only a coloured header bar would record them apart. Why the stylesheet holds back the first frame, and what preloading and inlining do about it, is the subject of What blocks the parser and the first paint, which also follows readyState, DOMContentLoaded and load through the same visit; the page-load card turns the same picture into LCP.
Who hands what to whom for the first frame (Chromium)
- Network service → Renderer main thread: HTML bytes, chunk by chunk
- Note over Renderer main thread: decode, tokenize, append DOM nodes
- Network service → Renderer main thread: main.css bytes
- Note over Renderer main thread: CSSOM, style, layout, pre-paint, paint
- Renderer main thread → Compositor thread: commit: display lists + property trees
- Note over Compositor thread: layerize, cut layers into tiles
- Compositor thread → Raster and decode threads: raster the tiles near the viewport first
- Raster and decode threads → Compositor thread (reply): tiles ready in GPU memory
- Compositor thread → GPU process (Viz): compositor frame: which tile goes where
- Note over GPU process (Viz): aggregate with other frames on screen, draw
In practice.
After the first frame the pipeline keeps running, but rarely in full. Each thing a Copperpot customer does wakes up a different part of it, and a good front-end design says which part on purpose. When those reruns happen (at most once per frame, between tasks) is the subject of The frame and the event loop.
One visit to the casserole page, stage by stage
Everything ran once: parse, style, layout, paint, raster, composite. The photo slot is laid out at its final size from the width and height attributes, so when the decoded photo arrives a frame later nothing below it moves.
First frame. Store “Copperpot”, showing the product view. Cart button “Basket” with badge 0 (its accessible name says “Cart, 0 items”, not the number alone). Product (showing) “Enamel casserole” by “Cast iron, oven safe to 260 °C”: rating 4.8 (“312 reviews”); price £64.00. Gallery: image 1 of 5 (“Cream casserole, lid on, side view”); thumbnails beside it. Colour (swatches, a radio group): Cream (selected), Teal, Rust. Size (tiles, a radio group): 24 cm (selected), 28 cm. Stock: “In stock”. Quantity 1 (stepper “Quantity”). Buttons: “Add to basket”. Delivery: “Free delivery over £75”. Folded sections: Details, Reviews, Care and cleaning. Note on gallery: box sized before the image decodes
- Its own layer while swiping3Compositor thread
- Text change means layout2Renderer main thread
- Boxes created only when opened2Renderer main thread
Which stages each interaction runs
- JavaScript, width 1, Runs
- Style, width 1, Runs
- Layout, width 1, Runs
- Paint, width 1, Runs
- Raster, width 1, Runs off the main thread
- Composite, width 1, Runs off the main thread
- Runs
- Skipped (output reused)
- Runs off the main thread
As it starts. 6 steps follow.
Main-thread work behind Copperpot's first frame
Data
| Stage | ms |
|---|---|
| Parse HTML | 56 |
| Parse CSS | 6 |
| Style | 14 |
| Layout | 14 |
| Paint | 7 |
| Raster (off main) | 8 |
The same interactions in Copperpot's code, with the route each one takes
/* style → paint → composite: colour never moves a box */
.swatch:hover { border-color: var(--rust); }
/* composite only, once the track has its own layer */
.gallery-track {
transition: transform 220ms ease-out;
}
/* the photo's box is reserved before the bytes arrive */
.gallery img { aspect-ratio: 1; width: 100%; height: auto; }
/* closed by default: no boxes, no layout cost */
.reviews[hidden] { display: none; }
// style → layout → paint → composite: text can change size
sizeTiles.addEventListener('change', (e) => {
price.textContent = prices[e.target.value];
stock.textContent = stockLine(e.target.value);
});
// composite only: no geometry is read or written
function showPhoto(i) {
track.style.transform = `translateX(${-i * 100}%)`;
}
// the expensive one: 40 cards join the layout tree
reviewsToggle.addEventListener('click', () => {
reviews.hidden = !reviews.hidden;
});
Seeing the stages for yourself
Every major browser's developer tools can record a trace of the main thread. In Chrome's Performance panel the stages in this topic appear under their own names, such as Parse HTML, Recalculate style, Layout and Paint, with the commit and the compositor's work on other tracks. Record Copperpot's size change and you will see a short recalculate style, a layout and a paint; record the gallery swipe and the main thread should be nearly empty while frames keep arriving. Paint flashing in Chrome's Rendering tab colours each repainted area green, which turns the hover and the swipe into an instant visual check of which route a change took.
A trimmed, annotated main-thread trace of Copperpot's first load and one size change
2 ms Parse HTML chunk 1 → head, header, title, price
66 ms Parse HTML chunk 2 → gallery, options
114 ms Parse HTML chunk 3 → reviews, footer; DOM complete
(idle: nothing to style until main.css arrives)
180 ms Parse stylesheet main.css → CSSOM
186 ms Recalculate style ~1,200 elements get computed styles
200 ms Layout every box gets a size and position
214 ms Paint display lists recorded, not pixels
221 ms Commit handed to the compositor thread
(raster tiles and the frame happen on other tracks)
Event: change handler sets two textContent values
Recalculate style two paragraphs
Layout the price, stock line and what follows
Paint the changed area
Commit
Where this shows up in system designs
Saying it in an interview
Trade-offs.
The pipeline's design is a set of trade-offs: when to show a first frame, and what to keep out of the layout tree. Each choice buys speed with something else: a blank screen, a jump, a gap or an accessibility trap. Skipping layout and paint for long off-screen sections with containment and content-visibility is covered in Rendering less.
- Pro:The first frame is styled, so nothing jumps from unstyled to styled
- Pro:Text and layout appear while the hero photo is still downloading, in a box already sized for it
- Pro:Later chunks of HTML are painted as they arrive, so long pages fill in from the top
- Con:A slow stylesheet keeps the whole screen blank, even when the DOM is complete
- Con:Images pop in later, and without reserved dimensions they push content down
A flash of unstyled content: raw HTML in default styles, then a jump to the real design; Layout and paint run twice for the same content, and anything the visitor was about to tap moves
The screen stays blank until the slowest image or frame finishes, 300 ms instead of 233 ms on this simplified page and far worse on a slow network; Text the visitor could already read is held back behind pixels they may not scroll to
- Pro:No boxes, so it costs nothing in layout or paint while closed
- Pro:Removed from the accessibility tree and tab order, matching what is on screen
- Con:Opening it creates its boxes and lays out everything after it
- Con:Its size cannot be measured while closed
Leaves a blank gap the size of the guide on the page; Still laid out on every layout
Still painted, still clickable and still read out by screen readers; Invisible controls can trap keyboard users