Render or otherwise correctly display HTML content in preview snippets #85

Closed
opened 2026-07-02 02:05:08 +00:00 by JMR-dev · 1 comment
JMR-dev commented 2026-07-02 02:05:08 +00:00 (Migrated from github.com)

Problem

snippetOf() (MailRepositoryImpl.kt:371-372) builds the mailbox list preview by stripping markup with
a single regex and collapsing whitespace:

private fun snippetOf(body: String): String =
    body.replace(Regex("<[^>]*>"), " ").replace(Regex("\\s+"), " ").trim().take(SNIPPET_LENGTH)

It runs unconditionally on both HTML and plain-text bodies (called from openMessage at
MailRepositoryImpl.kt:86 and prefetchMessage at :131) even though content.isHtml is available at
both call sites and never consulted. That causes two distinct problems:

  1. HTML bodies: the regex only strips tag delimiters, not element content, so a <style>/<script>
    block's raw CSS/JS text can leak straight into the snippet (the text between the two tags the regex
    removes survives). It also never decodes entities, so &nbsp;, &amp;, or numeric character
    references show up literally instead of the character they represent.
  2. Plain-text bodies: the same regex still runs, so a plain-text message containing literal </>
    (a quoted <user@example.com> header, a code snippet, "3 < 5") gets that text silently eaten as if it
    were markup, even though plain text needs no HTML handling at all.

Suggested fix

Branch on content.isHtml:

  • HTML: strip <script>…</script> / <style>…</style> blocks including their content first (e.g.
    Regex("(?is)<(script|style)[^>]*>.*?</\\1>")), then decode entities and normalize remaining markup —
    android.text.Html.fromHtml(html, Html.FROM_HTML_MODE_COMPACT).toString() handles entity decoding and
    reasonable block-level spacing with no new dependency (the reader's own HTML rendering uses a WebView
    via HtmlBody.kt, which is too heavyweight for a one-line snippet).
  • Plain text: just collapse whitespace and truncate — skip tag/entity handling entirely.
  • Keep the SNIPPET_LENGTH = 140 truncation as the last step either way.

Acceptance criteria

  • An HTML email with <style>/<script> blocks doesn't leak their text into the snippet.
  • An HTML email with entities (&amp;, &nbsp;, etc.) shows the decoded character in the snippet.
  • A plain-text message containing literal </> is left untouched aside from whitespace
    collapsing/truncation.
  • Unit test covering snippetOf (or its replacement) for HTML-with-style, HTML-with-entities, and
    plain-text-with-angle-brackets cases.
## Problem `snippetOf()` (`MailRepositoryImpl.kt:371-372`) builds the mailbox list preview by stripping markup with a single regex and collapsing whitespace: ```kotlin private fun snippetOf(body: String): String = body.replace(Regex("<[^>]*>"), " ").replace(Regex("\\s+"), " ").trim().take(SNIPPET_LENGTH) ``` It runs unconditionally on both HTML and plain-text bodies (called from `openMessage` at `MailRepositoryImpl.kt:86` and `prefetchMessage` at `:131`) even though `content.isHtml` is available at both call sites and never consulted. That causes two distinct problems: 1. **HTML bodies:** the regex only strips tag *delimiters*, not element content, so a `<style>`/`<script>` block's raw CSS/JS text can leak straight into the snippet (the text between the two tags the regex removes survives). It also never decodes entities, so `&nbsp;`, `&amp;`, or numeric character references show up literally instead of the character they represent. 2. **Plain-text bodies:** the same regex still runs, so a plain-text message containing literal `<`/`>` (a quoted `<user@example.com>` header, a code snippet, "3 < 5") gets that text silently eaten as if it were markup, even though plain text needs no HTML handling at all. ## Suggested fix Branch on `content.isHtml`: - **HTML:** strip `<script>…</script>` / `<style>…</style>` blocks *including their content* first (e.g. `Regex("(?is)<(script|style)[^>]*>.*?</\\1>")`), then decode entities and normalize remaining markup — `android.text.Html.fromHtml(html, Html.FROM_HTML_MODE_COMPACT).toString()` handles entity decoding and reasonable block-level spacing with no new dependency (the reader's own HTML rendering uses a WebView via `HtmlBody.kt`, which is too heavyweight for a one-line snippet). - **Plain text:** just collapse whitespace and truncate — skip tag/entity handling entirely. - Keep the `SNIPPET_LENGTH = 140` truncation as the last step either way. ## Acceptance criteria - [ ] An HTML email with `<style>`/`<script>` blocks doesn't leak their text into the snippet. - [ ] An HTML email with entities (`&amp;`, `&nbsp;`, etc.) shows the decoded character in the snippet. - [ ] A plain-text message containing literal `<`/`>` is left untouched aside from whitespace collapsing/truncation. - [ ] Unit test covering `snippetOf` (or its replacement) for HTML-with-style, HTML-with-entities, and plain-text-with-angle-brackets cases.
JMR-dev commented 2026-07-02 02:46:59 +00:00 (Migrated from github.com)

Checked against the two open PRs touching this area: PR #45 (screen-lock) doesn't touch this file at all; PR #46 (fetch-all + retention) changes MailRepositoryImpl.kt but only to thread uid through updateHeaderContent — snippetOf() is untouched. Still open after either merges.

Checked against the two open PRs touching this area: PR #45 (screen-lock) doesn't touch this file at all; PR #46 (fetch-all + retention) changes `MailRepositoryImpl.kt` but only to thread `uid` through `updateHeaderContent` — `snippetOf()` is untouched. Still open after either merges.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMail#85