feat(search): Unicode-aware case-insensitive search via casefold columns #234

Merged
JMR-dev merged 13 commits from feat-unicode-search-casefold into main 2026-07-03 23:29:04 +00:00
JMR-dev commented 2026-07-03 16:43:01 +00:00 (Migrated from github.com)

Restores Unicode-aware case-insensitive substring search (approach A from the #227 spike). Paging (#223) moved search to a SQL LIKE scan whose case-folding is ASCII-only, so non-ASCII terms stopped matching case-insensitively.

⚠️ Schema migration (v18→v19) — review requested, not auto-merged.

Change

  • Per-field casefold columns on messages: senderFold, senderEmailFold, subjectFold, snippetFold, each = lowercase(source) (Kotlin's lowercase() is Unicode-aware).
  • Why per-field, not one concatenated column: the searchable fields are maintained by partial UPDATEs that don't carry all four (updateBody fills the snippet post-insert; updateHeaderContent refreshes headers), so a single column couldn't stay in sync. Each fold depends only on its own source, kept current via thin DAO default-method wrappers — so all 5 call sites are unchanged.
  • Search matches the fold columns with a pattern built from the lowercased query.
  • Migration v18→v19 (additive columns + lower() backfill; the backfill is ASCII-only, so existing non-ASCII rows re-fold on their next write/sync).

Tests

  • MigrationTest.migrate18To19_addsAndBackfillsCasefoldSearchColumns — validates the columns + backfill against the exported v19 schema (runs on-device in CI).
  • MailRepositoryImplTest — asserts the query is casefolded ("ÄPFEL" → "%äpfel%"); the existing LIKE-escape test still holds.
  • Preflight green: build, unit tests, androidTest compile, lint, ktlint, detekt.

Closes #232

🤖 Generated with Claude Code

Restores **Unicode-aware case-insensitive substring search** (approach A from the #227 spike). Paging (#223) moved search to a SQL `LIKE` scan whose case-folding is ASCII-only, so non-ASCII terms stopped matching case-insensitively. > ⚠️ **Schema migration (v18→v19) — review requested, not auto-merged.** ## Change - **Per-field casefold columns** on `messages`: `senderFold`, `senderEmailFold`, `subjectFold`, `snippetFold`, each `= lowercase(source)` (Kotlin's `lowercase()` is Unicode-aware). - **Why per-field, not one concatenated column:** the searchable fields are maintained by *partial* `UPDATE`s that don't carry all four (`updateBody` fills the snippet post-insert; `updateHeaderContent` refreshes headers), so a single column couldn't stay in sync. Each fold depends only on its own source, kept current via thin **DAO default-method wrappers** — so all 5 call sites are unchanged. - Search matches the fold columns with a pattern built from the **lowercased** query. - Migration **v18→v19** (additive columns + `lower()` backfill; the backfill is ASCII-only, so existing non-ASCII rows re-fold on their next write/sync). ## Tests - `MigrationTest.migrate18To19_addsAndBackfillsCasefoldSearchColumns` — validates the columns + backfill against the exported v19 schema (runs on-device in CI). - `MailRepositoryImplTest` — asserts the query is casefolded (`"ÄPFEL"` → `"%äpfel%"`); the existing LIKE-escape test still holds. - Preflight green: build, unit tests, androidTest compile, lint, ktlint, detekt. Closes #232 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign in to join this conversation.