fix(mailbox): derive plain-text preview snippets from HTML bodies #106

Merged
JMR-dev merged 2 commits from fix-html-preview-snippets into main 2026-07-02 04:38:48 +00:00
JMR-dev commented 2026-07-02 03:03:47 +00:00 (Migrated from github.com)

Summary

snippetOf() built every mailbox preview snippet with a single tag-delimiter regex, ignoring the content.isHtml flag available at both of its call sites. This PR fixes both failure modes from the issue and backfills existing rows:

  • New Snippet.of(body, isHtml) (org.libremail.data.Snippet) replaces snippetOf() at both openMessage/prefetchMessage call sites in MailRepositoryImpl:
    • HTML bodies are reduced via the existing pure-Kotlin HtmlToText utility — <script>/<style> blocks are dropped including their content, tags stripped, entities decoded — then whitespace-collapsed. (No WebView, no Android framework classes; derivation stays JVM-unit-testable and runs once at fetch time, never per list row.)
    • Plain-text bodies get no markup handling at all, so literal <user@example.org> / 3 < 5 text survives; only whitespace collapsing applies.
    • The 140-char cap is kept as the last step on both paths.
  • HtmlToText entity decoding is now a single pass that also decodes decimal/hex numeric character references (&#8217; / &#x2019;), leaves unknown/malformed references as-is, and never re-decodes a produced character (&amp;lt; → &lt;).
  • Backfill: snippets are persisted when a body is first fetched and are never re-derived (sync prefetch only touches unfetched rows), so existing rows would keep their broken snippets forever. A data-only v13→v14 migration (MIGRATION_13_14, rebased on top of #46's v12→v13) re-derives the snippet of every row with a cached body using the corrected logic. No schema change relative to v13 — 14.json is identical to 13.json apart from the version/identity hash — and the exported schema is committed.

Acceptance criteria mapping

  • HTML <style>/<script> text no longer leaks → HtmlToText content-dropping + SnippetTest, repository wiring test
  • Entities (&amp;, &nbsp;, numeric refs) decode to their characters → single-pass decoder + tests
  • Plain text with literal </> untouched apart from collapsing/truncation → isHtml branch + tests
  • Unit tests cover HTML-with-style, HTML-with-entities, and plain-text-with-angle-brackets → SnippetTest (9 cases), HtmlToTextTest (+3 cases), MailRepositoryImplTest (+2 wiring tests asserting the persisted snippet)

Closes #85

🤖 Generated with Claude Code

## Summary `snippetOf()` built every mailbox preview snippet with a single tag-delimiter regex, ignoring the `content.isHtml` flag available at both of its call sites. This PR fixes both failure modes from the issue and backfills existing rows: - **New `Snippet.of(body, isHtml)`** (`org.libremail.data.Snippet`) replaces `snippetOf()` at both `openMessage`/`prefetchMessage` call sites in `MailRepositoryImpl`: - **HTML bodies** are reduced via the existing pure-Kotlin `HtmlToText` utility — `<script>`/`<style>` blocks are dropped *including their content*, tags stripped, entities decoded — then whitespace-collapsed. (No WebView, no Android framework classes; derivation stays JVM-unit-testable and runs once at fetch time, never per list row.) - **Plain-text bodies** get no markup handling at all, so literal `<user@example.org>` / `3 < 5` text survives; only whitespace collapsing applies. - The 140-char cap is kept as the last step on both paths. - **`HtmlToText` entity decoding** is now a single pass that also decodes decimal/hex numeric character references (`&#8217;` / `&#x2019;`), leaves unknown/malformed references as-is, and never re-decodes a produced character (`&amp;lt;` → `&lt;`). - **Backfill:** snippets are persisted when a body is first fetched and are never re-derived (sync prefetch only touches unfetched rows), so existing rows would keep their broken snippets forever. A **data-only v13→v14 migration** (`MIGRATION_13_14`, rebased on top of #46's v12→v13) re-derives the snippet of every row with a cached body using the corrected logic. No schema change relative to v13 — `14.json` is identical to `13.json` apart from the version/identity hash — and the exported schema is committed. ## Acceptance criteria mapping - HTML `<style>`/`<script>` text no longer leaks → `HtmlToText` content-dropping + `SnippetTest`, repository wiring test - Entities (`&amp;`, `&nbsp;`, numeric refs) decode to their characters → single-pass decoder + tests - Plain text with literal `<`/`>` untouched apart from collapsing/truncation → `isHtml` branch + tests - Unit tests cover HTML-with-style, HTML-with-entities, and plain-text-with-angle-brackets → `SnippetTest` (9 cases), `HtmlToTextTest` (+3 cases), `MailRepositoryImplTest` (+2 wiring tests asserting the persisted snippet) Closes #85 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign in to join this conversation.