Skip the ~1.5 MB catalog parse and the table scan on every launch — the
main startup cost. `seedBundledIfNeeded` re-seeds only when the recorded
version differs from `speciesCatalogVersion`, writing the version after
the seed commits so an interrupted seed retries next launch.
Adds `parseSpeciesCatalogVersion` and a test keeping the constant in sync
with the asset's `version` field.
Infer the catalog species a free-text variety label names ("Maiz de la
abuela" -> Zea mays) and prefill the category from the species family.
- Pure, testable matcher (domain/species_autoclassify.dart): whole-word,
accent/case-insensitive, Unicode-aware (any script), longest-name-wins,
ambiguous names left unclassified, light plural fold.
- Quick-add and draft naming auto-link the species when the field is empty
(non-destructive; an explicit category is kept).
- Edit sheet offers a one-tap suggestion from the typed name.
Tests: matcher unit, repository integration, SpeciesRepository.classifyLabel,
and an edit-sheet widget test.
Grow the bundled species catalog from 14 hand-curated entries to ~1200
edible/cultivated species, internationalized in 13 languages (es, en, fr,
de, it, pt, ca, gl, eu, ar, zh, ja, ru — Latin + Arabic RTL + CJK + Cyrillic).
- Add a reproducible generator (tool/gen_species_catalog.dart) that queries
Wikidata (CC0, no attribution burden) in two phases, filters out
non-vernacular noise (author citations, ranks, initials) and applies a
relevance floor, then merges hand-curated, authoritative core-crop data
(tool/curated_overrides.json: names, family, viability_years). GBIF is used
only as an identifier. The generated species.json (v3) is committed.
- Carry wikidata_qid and gbif_key through the parse/seed pipeline; the columns
already existed, so no DB migration.
- Rewrite seedBundled to one read + one batch (was a SELECT per species on
every startup — a real cost at ~1200 rows) and keep it idempotent with
backfill of the new reference fields.
- Move species search filtering to SQL (LIKE) so a large catalog is not pulled
into memory on every keystroke.
- Cover the generator transform, the generated asset, and the new fields with
tests.
Add a read-only "Learn more" section on the variety detail that surfaces
verified external references for the linked species, built deterministically
from bundled identifiers — no titles bundled, no network, offline-first:
- GBIF taxon page from the backbone key
- Wikipedia via Wikidata Special:GoToLinkedPage (locale-aware, RTL-safe,
falls back to the Wikidata item when a language lacks the article)
- Wikispecies from the scientific name
Enrich the bundled catalog (species.json v3) with gbif_key + wikidata_qid for
all 14 species, cross-checked (each QID carries GBIF ID P846 = the key).
Schema already had the columns; seedBundled now persists and backfills them.
The grower's own manual links stay but are demoted (progressive disclosure):
the section only shows when links exist and "Add link" is a low-emphasis
button, since hand-adding a link is the rarer power move.
Pure-Dart link builder is unit-tested; repo persistence/backfill and the
detail-screen section are covered too.
Lower the bulk-digitization cliff with two more routes on top of the
already-landed CSV import and "save and add another":
- Photo-first drafts (capture now, catalogue later): burst-capture
photos (camera or multi-gallery) into unnamed draft varieties, shown
in a "to catalogue" tray, hidden from the main list until named.
Adds Variety.isDraft (schema), addDraftVariety/watchDrafts/nameDraft,
the triage sheet and the inventory banner.
- On-device OCR label suggestion (Tesseract, offline, no Google): a
"Suggest name from photo" button in the naming dialog behind a
LabelTextExtractor interface (Tesseract on Android/iOS, no-op
elsewhere). Reads the largest print via hOCR bounding boxes, drops
boilerplate/low-confidence noise, preprocesses (grayscale, contrast,
upscale) and sweeps rotations (0-315 deg) so tilted packets still
read. Bundles tessdata_fast eng+spa; validated on-device against real
packets. The photo is written to a temp file deleted immediately in a
finally block (the plugin needs a path) - a bounded, documented
exception to no-plaintext-at-rest.
This commit also carries the co-developed schema evolution v5 to v8 that
shares these files (organic flag, species viability years, crop
calendar, lot provenance/abundance/preservation format, condition
checks) plus their exports/migrations and i18n.
Tests: CSV/draft/OCR unit + widget + migration green in isolation.
Note: the full widget suite currently hangs (>10 min) - under investigation.
Add a small curated catalog of Iberian horticultural species and let a variety
be linked to it from the edit sheet.
- assets/catalog/species.json: 14 species with botanical family and ES/EN
common names (wikidata_qid/gbif_key deferred to the varilla enrichment).
- SpeciesRepository: idempotent seedBundled (is_bundled rows, keyed by
scientific name) + search by scientific/common name with a locale-best label.
Seeded on startup from DI.
- VarietyRepository.linkSpecies: sets species_id and prefills category from the
species' family when empty (never overwrites an existing category).
VarietyDetail now carries the scientific name.
- Edit sheet gains a live species-search field; the detail view shows the
scientific name (italic). i18n strings added (ES/EN).
Tests: catalog parse, idempotent seeding, search by scientific/common name,
linkSpecies prefill semantics, and a widget test for the autocomplete → link →
scientific-name-shown flow. Full suite: 32 passing, 0 skipped.