What is HTML size?
Every page starts as one file of HTML code. The browser has to download and read that file before it can show anything or fetch the images, styles and scripts the page needs. The bigger the file, the later everything else starts.
There are two sizes worth knowing. The first is the size of the code as written. The second is the size that actually travels over the network: most servers compress HTML before sending it, which often shrinks it to a fifth of its size or less.
How this test works
Our server downloads your page the way a browser first receives it and reads the code without running any scripts. It measures the whole file, then sorts the code into kinds: tags and text, scripts and styles written into the page, SVG images and comments.
The chart “Size by tag and attribute” goes further and files every piece of the page under the tag and attribute it belongs to. “img · src” is the total size of every image address on the page; “div · class” is all the class names on all the div tags. An entry ending in “text” is the text directly inside that tag, plus the tag itself without its attributes, so “script · text” is the code written inside script tags. An SVG image is counted whole, as one item. This is where bloat shows up: images embedded as text appear under “img · src”, and the data a JavaScript framework ships with the page appears under “script · text”.
Each row shows three figures: the item’s share of the page, its full size, and its size after compression with the method your site uses. Repeated code, such as the same class names on hundreds of tags, shrinks a lot when compressed. Embedded images hardly shrink at all.
Why the list is sorted by full size
We sort by the size before compression, for three reasons. First, compression only helps the trip over the network: once the page arrives, the browser unpacks it and has to read every byte of the full code, so the full size is what decides how long that takes. Second, Google counts its 2 MB limit on the uncompressed page. Third, the full size is exact, while the compressed figure for a single item is only a guide. We compress each item on its own, and a whole page compresses better than its parts, so the items add up to more than the compressed page.
If your site uses Brotli or Zstandard and our server can’t run that method, the compressed figures are worked out with gzip instead, and the chart says so.
The row of figures at the top of the report compares compression methods. We ask your server for the page three more times, each time saying we accept only one method: gzip, the long-standing standard; Brotli, the method most sites use today; and Zstandard, a newer one that recent browsers also support. We measure each reply exactly as it arrives. If your server answers without the method we asked for, its figure is replaced by “not offered”. These are your server’s real results, with its own settings, not our calculation. One figure carries a tick: that is the method your server chose when we offered all of them at once, the way a browser does, so it is what your visitors are getting now. If we can’t measure any of them, we compress the page ourselves with gzip and mark the figure as an estimate.
What can be trimmed from the code
Some characters in a page’s code do nothing for the browser, and a minifier can remove them without changing how the page looks or works. The report estimates five such savings, the same ones the open-source PageSpeed optimisation module makes:
- extra spaces and line breaks, where a run of them can become one;
- comments left for developers;
- quotes around attribute values that don’t need them, such as class="menu";
- attributes that only repeat what a browser assumes anyway, such as type="text/javascript" on a script;
- your site’s own address at the start of links to your own pages, where a shorter relative link works the same.
The table shows, for each one, the bytes that would go from the code as written and roughly what that saves once the page is compressed. The second number is much smaller, because repeated characters such as spaces already compress very well. That is worth knowing before you spend time on it: on a site with compression turned on, minifying HTML is a small gain.
We don’t count spacing inside scripts, styles and preformatted text, where it can matter, or the canonical and hreflang links, which are meant to be written in full. Removing comments can break a tool that reads them, and shortening links can break a page that is copied to another address, so test after turning any of this on.
What can be moved out of the page
JavaScript, CSS and SVG images written directly into a page can almost always be moved into files of their own. The page then only points to them. The gain is that a browser keeps those files after the first visit and reuses them on every other page, instead of downloading the same code again inside each page. The report shows how much of each your page carries, as written and after compression.
A few things are better left in the page: structured data for search engines, which has to be in the page to be read, and the small amount of CSS needed to draw the top of the page straight away.
The report also counts the page’s DOM: elements are the tags themselves, and nodes are everything in the tree, including the pieces of text and the comments between tags. A large DOM takes a browser longer to build and to update.
What is a good result?
There is no official target, so the score is built on our own marks. It looks at two sizes. For the code as written, 75 KB or less scores 100 and 250 KB or more scores 0. For the size sent over the network after compression, 20 KB or less scores 100 and 50 KB or more scores 0. Between those marks each falls in a straight line. The report shows the two scores side by side, each in its own ring, so you can see at once whether the weight is in the code itself or in how it is sent. The heading above them goes by the lower one. A page sent without compression is judged on its full size for both, because that is what your visitors download.
The cards in the report explain the score but don’t change it: they show what takes up the space and what to do about it.
For a page to count as lean, we aim higher: 50 to 75 KB of HTML as written, which is 15 to 20 KB once compressed, with readable text making up at least 25% of the code. A page that size is quick for browsers, search engine crawlers and AI crawlers alike to download and read in full. This target is Seokla’s own guide, drawn from our experience; no search engine or AI company publishes such a figure. The report estimates how close your page would get if its inline JavaScript, CSS and SVG images were moved into separate files.
One limit is official: Google says Googlebot reads only the first 2 MB of a page’s HTML, counted before compression. Anything after that point is not considered for indexing. The report tells you how close your page is.
Should a page follow the HTML standard?
Yes, and it is worth doing, though not for the reason often given. The HTML standard, maintained by the WHATWG, describes how a page’s code should be written. Google has said more than once that valid HTML is not a ranking factor, and many well-known sites don’t pass a validator, because browsers are built to cope with imperfect code.
Errors can still cost you. Google’s own documentation says that when it meets an element that doesn’t belong in the head of a page, such as an image or an iframe, it treats the head as finished and stops reading. Anything after that point can be missed, including a canonical link, hreflang tags or instructions for robots. Outside the head, a browser has to guess what broken code was meant to be, and its guess isn’t always what you intended. Screen readers also depend on correct structure to describe a page to people who can’t see it.
There is a link to size as well. Some of what bloats a page, such as tags used only for styling and layers of wrappers that carry no meaning, is exactly what the standard steers authors away from.
This tool measures size only. It doesn’t check your code against the standard. For that you need an HTML validator, a tool that reads the code and lists where it departs from the standard. Treat its report as a list of things to look at, not a score to chase: fix what breaks the head, the structure or accessibility first.
What this test can’t tell you
This is the size of the HTML file alone. Images, stylesheets, fonts and script files are downloaded separately and are usually much larger; this test doesn’t count them. The download times are estimates worked out from the compressed size and an assumed connection speed, not measurements. We check pages up to 3 MB, so a larger page is reported as “over 3 MB”.