# GerbilManager import tooling (FEAT-8b) One-off **migration tooling** (Python, no third-party deps) that turns Julian's wife's hand-built spreadsheets into normalised JSON for review and, later, import. This is *not* product code — it lives outside the app and is run manually. See the format analysis in `FEAT-8a-format-spec.md` (Pam's hive workspace). ## What it does `extract.py` runs **stages 1–2** of the pipeline: 1. **Extract (stage 1)** - 10 *Stammbaum* pedigree charts → animals (name, DOB, death, Farbschlag, genotype, breeder, positionally-reconstructed parent links, photos). - *Wurfchronik* litter chronicle → litters (date, dam, sire, Wurfstärke, sex breakdown, Zuchtnummer, notes). Columns are read **by header row** because the two sheets use different schemas. - Embedded photos (`xl/media`) → `output/photos//`, mapped to the animal by drawing anchor position. 2. **Dedup + review (stage 2)** - Merge animals on `normalise(call-name) + DOB`, with the **Zucht as discriminator** (Julian's ruling: Wurfchronik `[brackets]` ≡ Stammbaum `of/von ` suffix — both are the breeding line; same name+DOB but different Zucht stays two animals). - Match animals onto Wurfchronik litters (`litterRef`) via DOB + (Vater, Mutter) — the Pam-validated build order (chronicle litters are canonical). - Emit a German-language `output/review-report.md` for the breeder to verify (merges, **conflicts**, ambiguous/incomplete entries, unmapped genotype tokens, litter data-quality warnings). - **Nothing is loaded into the database** — stage 3 (API load) is separate and waits on DATA-2 + FEAT-1b phase 2. ### Wurfchronik column semantics (Julian, authoritative) `A` Wurfbezeichnung · `B` Geburtsdatum · `C` Mutter · `D` Vater (`[…]` = Zucht, `&` = multiple sires) · `E` **survivedToGoHome** (Tabelle1 only, unlabeled — detected positionally) · `F` Wurfstärke → `totalBorn` · `G` breakdown `Männchen,Weibchen,TG,s` → `males/females/stillborn/diedLater` (`s` = died after birth, before Abgabe) · last column → `note`. Validation: `E` should equal `F − TG − s`; mismatches become German warnings in the review report (data-quality signal, not an import blocker). A few Tabelle2 rows shift these columns — they are read value-adaptively and flagged with a warning. Genotypes are mapped to the frozen 8-locus contract (A C D E G P Sp Re) while preserving everything: `genotype.mapped8locus`, `genotype.rawGenotype` (verbatim), `genotype.unmappedTokens` (e.g. the `Uw` locus, markers `WFNZ/WP/DP`). A `-` (unknown second allele) maps to `?`. ## Run ```sh cd tools/import python extract.py # xlsx → animals.json / litters.json python extract.py --stammbaeume "" --wurfchronik "" python extract_docx.py # Wurfchronik-Detail.docx → docx_*.json python extract_contracts.py # Abgabeverträge (.docx) → contracts.json python merge_and_resolve.py # → resolved_import.json (DB-ready) ``` Requires Python 3 (zero third-party deps). **Re-runnable / idempotent** — re-run when more files arrive (Wurfchronik `Teil2+`, new charts, or new contracts). `extract_contracts.py` scans the breeder's sale-contract share (`\\truenas\…\Verträge`, ~1.4k `.docx`) and emits one record per contract (buyer, animal call-names, Farbschlag, dates, price, source filename). It skips the blank template, `Abstammungsnachweis`/`Geburtsurkunde` documents, and any file that is not a readable `.docx`. `merge_and_resolve.py` then conservatively folds contracts into the resolved data: buyers become receiver `Contacts`, and unambiguously matched gerbils get `ReceiverContactId` / `GoHomeDate` / `Status=GivenAway` (only where not already set), with a provenance history line. Ambiguous / unmatched animals are counted and skipped, never guessed. ## Output (`tools/import/output/`, git-ignored except the report) | File | Contents | |---|---| | `animals.json` | deduped animals with genotype, parentRefs, photos, sourceFiles | | `litters.json` | litters from the Wurfchronik | | `docx_animals.json` / `docx_litters.json` | Wurfchronik-Detail.docx rows | | `contracts.json` | one record per Abgabevertrag (buyer, animals, dates, price) | | `resolved_import.json` | merged DB-ready payload consumed by `IngestResolvedService` | | `photos//…` | extracted, anchor-mapped images | | `review-report.md` | **human review deliverable** (committed) | ## Files - `xlsx_util.py` — dependency-free `.xlsx` reader (zip + XML): shared strings, cells by reference, image/drawing anchors. - `genotype.py` — genotype notation parser → 8-locus mapping + raw + unmapped. - `extract.py` — the pipeline (stages 1–2).