tane/apps/app_seeds/integration_test/ocr_label_test.dart
vjrj 6809dc6143 feat(inventory): photo-first drafts + on-device OCR (digitization R2+R4)
Lower the bulk-digitization cliff with two more routes on top of the
already-landed CSV import and "save and add another":

- Photo-first drafts (capture now, catalogue later): burst-capture
  photos (camera or multi-gallery) into unnamed draft varieties, shown
  in a "to catalogue" tray, hidden from the main list until named.
  Adds Variety.isDraft (schema), addDraftVariety/watchDrafts/nameDraft,
  the triage sheet and the inventory banner.
- On-device OCR label suggestion (Tesseract, offline, no Google): a
  "Suggest name from photo" button in the naming dialog behind a
  LabelTextExtractor interface (Tesseract on Android/iOS, no-op
  elsewhere). Reads the largest print via hOCR bounding boxes, drops
  boilerplate/low-confidence noise, preprocesses (grayscale, contrast,
  upscale) and sweeps rotations (0-315 deg) so tilted packets still
  read. Bundles tessdata_fast eng+spa; validated on-device against real
  packets. The photo is written to a temp file deleted immediately in a
  finally block (the plugin needs a path) - a bounded, documented
  exception to no-plaintext-at-rest.

This commit also carries the co-developed schema evolution v5 to v8 that
shares these files (organic flag, species viability years, crop
calendar, lot provenance/abundance/preservation format, condition
checks) plus their exports/migrations and i18n.

Tests: CSV/draft/OCR unit + widget + migration green in isolation.
Note: the full widget suite currently hangs (>10 min) - under investigation.
2026-07-09 21:23:46 +02:00

38 lines
1.7 KiB
Dart

import 'package:flutter/services.dart' show rootBundle;
import 'package:flutter_test/flutter_test.dart';
import 'package:integration_test/integration_test.dart';
import 'package:tane/services/ocr/tesseract_label_extractor.dart';
/// Real, on-device OCR validation against actual packet photos — the "does
/// Tesseract really read my packets" gate from the plan. Runs the native engine,
/// so it must run on a device/emulator, NOT under `flutter test` on the host:
///
/// flutter test integration_test/ocr_label_test.dart -d <android-device>
///
/// Drop the photos to check into `assets/ocr_fixtures/` (declared in pubspec)
/// with the file names below. Each test prints what Tesseract actually read so
/// you can see and tune it, then asserts the expected variety word is present.
void main() {
IntegrationTestWidgetsFlutterBinding.ensureInitialized();
const extractor = TesseractLabelExtractor();
Future<void> reportRead(String asset) async {
final data = await rootBundle.load('assets/ocr_fixtures/$asset');
final result = await extractor.suggestLabel(data.buffer.asUint8List());
// ignore: avoid_print
print('OCR[$asset] => "$result"');
// Best-effort: a usable, editable suggestion should appear. Exact spelling
// is not asserted — real OCR may misread a letter (e.g. "Sucumber"), which
// is still useful as a prefill the user corrects.
expect(result, isNotNull, reason: 'Tesseract produced no suggestion for '
'$asset (see the printed value above)');
}
testWidgets('suggests a name from the AUBERGINE packet photo', (_) async {
await reportRead('aubergine.jpg');
});
testWidgets('suggests a name from the CUCUMBER packet photo', (_) async {
await reportRead('cucumber.jpg');
});
}