html_to_text()

Derive plain text from html.

Usage

Source

html_to_text(html)

It writes a link as label <url>, or as the label alone when the label is already the URL or the mailto: address, or when the URL is a fragment such as #top. It keeps the line breaks and indentation of a <pre> block, such as a code block or a log. It starts a list item with -. It writes a table row on one line with | between its cells when the cells hold only inline content. A block or a <br> inside a cell starts a new line, so the paragraphs of a layout table stay apart. It writes empty cells too, so it never shifts a value into the wrong column. It writes an image as its alt text in brackets. It drops <style>, <script>, <title>, and comments. It decodes entities. It returns "" for HTML that holds no text, such as a lone image with no alt text.

It never raises on malformed HTML. The output is best effort and not a contract, so it may change in a minor version. Pass text= to Message when the exact text matters. See ADR-0008.

Examples

from epistole import html_to_text

html = '<p>The report is <a href="https://example.com/q3">online</a>.</p>'
html_to_text(html)  # "The report is online <https://example.com/q3>."