fix(mailbox): derive plain-text preview snippets from HTML bodies #106

Merged
JMR-dev merged 2 commits from fix-html-preview-snippets into main 2026-07-02 04:38:48 +00:00
2 Commits
Author SHA1 Message Date
Jason Ross c18ab42c6c Merge branch 'main' into fix-html-preview-snippets 2026-07-01 23:29:44 -05:00
JMR-devandClaude Fable 5 c06a387b3c fix(mailbox): derive plain-text preview snippets from HTML bodies
snippetOf() stripped only tag delimiters with a single regex on every
body, HTML or not: <style>/<script> text leaked into HTML snippets,
entities stayed encoded, and plain-text bodies had literal <...> text
eaten as if it were markup.

Replace it with Snippet.of(body, isHtml), which finally consults the
isHtml flag both call sites already had: HTML bodies go through
HtmlToText (script/style content dropped, tags stripped, entities
decoded), plain text gets no markup handling at all; both paths keep
the whitespace collapsing and the 140-char cap. HtmlToText's entity
decoding is now a single-pass decoder that also handles decimal/hex
numeric character references and never re-decodes produced characters.

Snippets are persisted when a body is first fetched and never
re-derived, so existing rows would keep their broken snippets forever;
a data-only v13->v14 migration re-derives every cached row's snippet
with the corrected logic (schema unchanged relative to v13, exported
14.json committed).

Closes #85

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-01 23:13:16 -05:00