Files
GerbilManager/tools/import/README.md
Gulum 7ea2f97e72 FEAT-8b: bake Julian's authoritative Wurfchronik semantics into the extractor
- survivedToGoHome: unlabeled Tabelle1 col E detected positionally; value-
  adaptive row parsing recovers it from schema-shifted Tabelle2 rows too
  (118 recovered, 79 confirmed by the E=F-TG-s identity)
- breakdown G -> males/females/stillborn/diedLater ('s' = died before Abgabe)
- validation: E = F - TG - s; 113 mismatches as German review-report warnings
- (name, Zucht) canonicalisation: [brackets] == of/von suffix ([ZdkC] ==
  von den Kleinen Chaoten); Zucht = dedup discriminator (0 splits in data)
- animal->litter matching via DOB+(Vater,Mutter): 95 high-confidence,
  40 date-only, 9 ambiguous; litterRef in animals.json
- regenerated report: 889 raw -> 574 unique (279 dated), 32 conflicts

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 00:56:14 +02:00

76 lines
3.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GerbilManager import tooling (FEAT-8b)
One-off **migration tooling** (Python, no third-party deps) that turns Julian's
wife's hand-built spreadsheets into normalised JSON for review and, later, import.
This is *not* product code — it lives outside the app and is run manually.
See the format analysis in `FEAT-8a-format-spec.md` (Pam's hive workspace).
## What it does
`extract.py` runs **stages 12** of the pipeline:
1. **Extract (stage 1)**
- 10 *Stammbaum* pedigree charts → animals (name, DOB, death, Farbschlag,
genotype, breeder, positionally-reconstructed parent links, photos).
- *Wurfchronik* litter chronicle → litters (date, dam, sire, Wurfstärke,
sex breakdown, Zuchtnummer, notes). Columns are read **by header row** because
the two sheets use different schemas.
- Embedded photos (`xl/media`) → `output/photos/<animal-slug>/`, mapped to the
animal by drawing anchor position.
2. **Dedup + review (stage 2)**
- Merge animals on `normalise(call-name) + DOB`, with the **Zucht as
discriminator** (Julian's ruling: Wurfchronik `[brackets]` ≡ Stammbaum
`of/von <line>` suffix — both are the breeding line; same name+DOB but
different Zucht stays two animals).
- Match animals onto Wurfchronik litters (`litterRef`) via DOB + (Vater,
Mutter) — the Pam-validated build order (chronicle litters are canonical).
- Emit a German-language `output/review-report.md` for the breeder to verify
(merges, **conflicts**, ambiguous/incomplete entries, unmapped genotype
tokens, litter data-quality warnings).
- **Nothing is loaded into the database** — stage 3 (API load) is separate and
waits on DATA-2 + FEAT-1b phase 2.
### Wurfchronik column semantics (Julian, authoritative)
`A` Wurfbezeichnung · `B` Geburtsdatum · `C` Mutter · `D` Vater (`[…]` = Zucht,
`&` = multiple sires) · `E` **survivedToGoHome** (Tabelle1 only, unlabeled —
detected positionally) · `F` Wurfstärke → `totalBorn` · `G` breakdown
`Männchen,Weibchen,TG,s``males/females/stillborn/diedLater` (`s` = died
after birth, before Abgabe) · last column → `note`. Validation: `E` should
equal `F TG s`; mismatches become German warnings in the review report
(data-quality signal, not an import blocker). A few Tabelle2 rows shift these
columns — they are read value-adaptively and flagged with a warning.
Genotypes are mapped to the frozen 8-locus contract (A C D E G P Sp Re) while
preserving everything: `genotype.mapped8locus`, `genotype.rawGenotype` (verbatim),
`genotype.unmappedTokens` (e.g. the `Uw` locus, markers `WFNZ/WP/DP`). A `-`
(unknown second allele) maps to `?`.
## Run
```sh
cd tools/import
python extract.py # uses the default source paths
python extract.py --stammbaeume "<dir>" --wurfchronik "<file.xlsx>"
```
Requires Python 3. **Re-runnable / idempotent** — re-run when more files arrive
(Wurfchronik `Teil2+`, or new charts).
## Output (`tools/import/output/`, git-ignored except the report)
| File | Contents |
|---|---|
| `animals.json` | deduped animals with genotype, parentRefs, photos, sourceFiles |
| `litters.json` | litters from the Wurfchronik |
| `photos/<slug>/…` | extracted, anchor-mapped images |
| `review-report.md` | **human review deliverable** (committed) |
## Files
- `xlsx_util.py` — dependency-free `.xlsx` reader (zip + XML): shared strings,
cells by reference, image/drawing anchors.
- `genotype.py` — genotype notation parser → 8-locus mapping + raw + unmapped.
- `extract.py` — the pipeline (stages 12).