Python html.parser: Parse HTML

Published on: August 20, 2026
Reading time: 5 minutes
HTML code on a screen representing parsing with Python html.parser

The html.parser module provides an event-driven HTML and XHTML parser. You subclass HTMLParser and override methods that are called for start tags, end tags, text, comments, character references, and declarations.

It is lightweight, included in the standard library, and tolerant of many malformed documents found on the web. It works well for extracting links, titles, text, and metadata, building narrowly scoped validators, and processing HTML incrementally. It does not build a complete browser DOM, execute JavaScript, or sanitize untrusted markup.

How HTMLParser works

The parser accepts text through feed(). Complete elements trigger handler methods, while incomplete data remains buffered until more input arrives or close() is called.

from html.parser import HTMLParser

class DebugParser(HTMLParser):
    def handle_starttag(self, tag, attrs):
        print("start", tag, attrs)

    def handle_endtag(self, tag):
        print("end", tag)

    def handle_data(self, data):
        print("text", repr(data))

parser = DebugParser()
parser.feed("<h1>Hello &amp; Python</h1>")
parser.close()

With the default convert_charrefs=True, named and numeric references are converted to Unicode characters, except in special raw-text contexts such as script and style.

from html.parser import HTMLParser
from urllib.parse import urljoin

class LinkParser(HTMLParser):
    def __init__(self, base_url: str):
        super().__init__()
        self.base_url = base_url
        self.links: list[str] = []

    def handle_starttag(self, tag, attrs):
        if tag != "a":
            return

        attributes = dict(attrs)
        href = attributes.get("href")
        if href:
            self.links.append(urljoin(self.base_url, href))

parser = LinkParser("https://example.com/docs/")
parser.feed('<a href="python.html">Python</a>')
parser.close()
print(parser.links)

Tag and attribute names are lowercased, quotes are removed from values, and valueless attributes receive None. The urllib.parse guide explains URL resolution and validation.

Validate extracted URLs

A link can use javascript:, data:, file:, or an unexpected origin. The parser only returns a string; it does not decide whether it is safe.

from urllib.parse import urlsplit

parts = urlsplit(candidate)
if parts.scheme not in {"http", "https"}:
    return
if parts.hostname not in {"example.com", "www.example.com"}:
    return

When fetching discovered links, apply SSRF protection, redirect validation, size limits, and timeouts as described in the urllib.request guide.

Extracting readable text

handle_data() receives ordinary text, but also the content of script and style. Track context when generating readable text.

class TextParser(HTMLParser):
    ignored = {"script", "style", "noscript"}

    def __init__(self, **kwargs):
        super().__init__(**kwargs)
        self.skip_depth = 0
        self.parts: list[str] = []

    def handle_starttag(self, tag, attrs):
        if tag in self.ignored:
            self.skip_depth += 1

    def handle_endtag(self, tag):
        if tag in self.ignored and self.skip_depth:
            self.skip_depth -= 1

    def handle_data(self, data):
        if self.skip_depth == 0:
            cleaned = " ".join(data.split())
            if cleaned:
                self.parts.append(cleaned)

    def text(self) -> str:
        return " ".join(self.parts)

Malformed HTML may close elements in unexpected order. A simple counter is suitable for controlled inputs but does not replace a full HTML5 tree-construction algorithm.

Extracting title and meta description

class MetadataParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_title = False
        self.title_parts = []
        self.description = None

    def handle_starttag(self, tag, attrs):
        attributes = dict(attrs)
        if tag == "title":
            self.in_title = True
        elif tag == "meta" and attributes.get("name", "").lower() == "description":
            self.description = attributes.get("content")

    def handle_endtag(self, tag):
        if tag == "title":
            self.in_title = False

    def handle_data(self, data):
        if self.in_title:
            self.title_parts.append(data)

    @property
    def title(self):
        return " ".join("".join(self.title_parts).split())

Pages can contain duplicate tags, empty values, and metadata produced only by JavaScript. Define a selection policy and cap all accumulated content.

Incremental parsing

feed() accepts fragments and buffers incomplete tags.

parser = LinkParser("https://example.com/")

for chunk in ["<a hr", 'ef="/docs">Doc', "umentation</a>"]:
    parser.feed(chunk)

parser.close()

This fits streaming downloads and avoids storing the entire page. However, feed() requires str. Use an incremental decoder so a UTF-8 sequence is not split incorrectly. The Python codecs guide covers this boundary.

Incremental decoding with limits

import codecs

parser = TextParser()
decoder = codecs.getincrementaldecoder("utf-8")(errors="replace")
total = 0

while chunk := response.read(64 * 1024):
    total += len(chunk)
    if total > 5 * 1024 * 1024:
        raise ValueError("HTML is too large")
    parser.feed(decoder.decode(chunk))

parser.feed(decoder.decode(b"", final=True))
parser.close()

Determine encoding from a trustworthy policy. HTTP headers, BOMs, and HTML declarations may disagree. Complex scraping may require a library implementing the complete HTML encoding-sniffing rules.

Comments and declarations

handle_comment() receives comment contents without delimiters. handle_decl() receives declarations such as DOCTYPE.

class AuditParser(HTMLParser):
    def handle_comment(self, data):
        if "TODO" in data:
            print("review comment found")

    def handle_decl(self, decl):
        print("declaration", decl)

Comments are delivered to every client. Never put secrets in them.

Character references

With convert_charrefs=True, named and numeric references are converted automatically. To inspect the original lexical form, pass convert_charrefs=False and implement handle_entityref() and handle_charref().

Entity conversion is not sanitization. A string containing decoded angle brackets can become dangerous if it is later inserted as markup without contextual escaping.

The scripting parameter

Since Python 3.14.1, HTMLParser accepts scripting. When true, content inside noscript is returned as raw text rather than parsed as markup. This approximates parsing when scripting is enabled, but it does not execute JavaScript.

parser = TextParser(scripting=True)

If a subclass defines __init__, accept and forward required keyword arguments to super().__init__().

Malformed HTML and parser limits

The parser accepts invalid markup, but it does not check that end tags match start tags and does not synthesize all implicit closing events a browser tree builder would create. An event sequence may not represent the DOM produced by a browser.

For CSS selectors, structural editing, and HTML5 fidelity, choose Beautiful Soup, lxml, or html5lib. The Beautiful Soup scraping guide demonstrates a tree-based approach.

Parsing is not sanitization

HTMLParser does not remove scripts, event-handler attributes, dangerous URLs, or malicious CSS. Do not use it alone to allow user HTML on a website.

Sanitization requires a maintained library with explicit allowlists for tags, attributes, and protocols, followed by correct contextual escaping at render time.

Never execute extracted content

Do not pass script text to eval(), a shell, or another interpreter. Do not automatically open discovered URLs or use attribute values as filesystem paths. Treat all external HTML as untrusted.

Resource protection

Cap downloaded bytes, link count, individual attribute lengths, and accumulated text. A document can contain millions of tags or a single huge attribute.

if len(self.links) > 10_000:
    raise ValueError("Too many links")
if href and len(href) > 4_096:
    return

The parser does not provide a global resource policy; the subclass must enforce one.

Reuse and reset

Create one parser per document or call reset() and clear all subclass state as well. Forgetting lists and flags mixes results from separate pages.

Testing

Test unquoted attributes, references, comments, tags divided between chunks, incorrect nesting, scripts, relative URLs, empty values, and resource limits. Ensure close() is always called.

Common mistakes

Typical failures include passing bytes instead of text, forgetting close(), assuming browser-equivalent structure, collecting script text, accepting every URL, omitting limits, reusing stale state, confusing entity conversion with sanitization, and rendering extracted text as HTML.

Best practices

Use one instance per document, an incremental decoder, byte and event limits, URL validation, and explicit script handling. Choose a tree parser for structural work. Use a dedicated sanitizer and contextual escaping for user content.

Conclusion

html.parser is a lightweight event-driven extraction tool. It can process HTML incrementally, tolerate imperfect markup, and collect text, links, and metadata without dependencies. Its boundaries are important: it is not a browser, a complete HTML5 parser, or a security sanitizer.

Read the official html.parser documentation and the HTML Living Standard.

Share:

Facebook
WhatsApp
Twitter
LinkedIn

Article content

    Related articles

    Server rack representing an endpoint built with Python xmlrpc.server
    Advanced Python
    Foto de perfil de Leandro Hirt da Academify

    xmlrpc.server: Build XML-RPC Servers

    Learn Python xmlrpc.server to build XML-RPC servers, register functions, restrict methods and paths, and avoid unsafe exposure.

    Ler mais

    Tempo de leitura: 5 minutos
    21/08/2026
    Server cables representing remote calls with Python xmlrpc.client
    Advanced Python
    Foto de perfil de Leandro Hirt da Academify

    Python xmlrpc.client: Remote Calls

    Learn Python xmlrpc.client to call XML-RPC services, handle Fault and ProtocolError, use TLS, compatible types, and safe limits.

    Ler mais

    Tempo de leitura: 5 minutos
    21/08/2026
    Error code over binary data representing failures handled with Python urllib.error
    Advanced Python
    Foto de perfil de Leandro Hirt da Academify

    Python urllib.error: Handle HTTP Failures

    Learn Python urllib.error to handle URLError, HTTPError, incomplete downloads, selective retries, and clearer network diagnostics.

    Ler mais

    Tempo de leitura: 5 minutos
    21/08/2026
    HTML keycaps representing HTML entities with Python html.entities
    Advanced Python
    Foto de perfil de Leandro Hirt da Academify

    html.entities: Convert HTML Entities

    Learn Python html.entities to inspect HTML entities, convert names and code points, and avoid confusing decoding with sanitization.

    Ler mais

    Tempo de leitura: 6 minutos
    21/08/2026
    Folder with files representing MIME types identified with Python mimetypes
    Advanced Python
    Foto de perfil de Leandro Hirt da Academify

    mimetypes: Detect MIME Types for Files

    Learn Python mimetypes to identify file types, validate uploads, and set safer HTTP Content-Type headers.

    Ler mais

    Tempo de leitura: 6 minutos
    20/08/2026
    Server rack representing low-level HTTP connections with Python http.client
    Advanced Python
    Foto de perfil de Leandro Hirt da Academify

    Python http.client: Low-Level HTTP

    Learn Python http.client for low-level HTTP and HTTPS connections, streaming, headers, TLS, connection reuse, size limits, and errors.

    Ler mais

    Tempo de leitura: 4 minutos
    20/08/2026