- survivedToGoHome: unlabeled Tabelle1 col E detected positionally; value-
adaptive row parsing recovers it from schema-shifted Tabelle2 rows too
(118 recovered, 79 confirmed by the E=F-TG-s identity)
- breakdown G -> males/females/stillborn/diedLater ('s' = died before Abgabe)
- validation: E = F - TG - s; 113 mismatches as German review-report warnings
- (name, Zucht) canonicalisation: [brackets] == of/von suffix ([ZdkC] ==
von den Kleinen Chaoten); Zucht = dedup discriminator (0 splits in data)
- animal->litter matching via DOB+(Vater,Mutter): 95 high-confidence,
40 date-only, 9 ambiguous; litterRef in animals.json
- regenerated report: 889 raw -> 574 unique (279 dated), 32 conflicts
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
76 lines
3.5 KiB
Markdown
76 lines
3.5 KiB
Markdown
# GerbilManager import tooling (FEAT-8b)
|
||
|
||
One-off **migration tooling** (Python, no third-party deps) that turns Julian's
|
||
wife's hand-built spreadsheets into normalised JSON for review and, later, import.
|
||
This is *not* product code — it lives outside the app and is run manually.
|
||
|
||
See the format analysis in `FEAT-8a-format-spec.md` (Pam's hive workspace).
|
||
|
||
## What it does
|
||
|
||
`extract.py` runs **stages 1–2** of the pipeline:
|
||
|
||
1. **Extract (stage 1)**
|
||
- 10 *Stammbaum* pedigree charts → animals (name, DOB, death, Farbschlag,
|
||
genotype, breeder, positionally-reconstructed parent links, photos).
|
||
- *Wurfchronik* litter chronicle → litters (date, dam, sire, Wurfstärke,
|
||
sex breakdown, Zuchtnummer, notes). Columns are read **by header row** because
|
||
the two sheets use different schemas.
|
||
- Embedded photos (`xl/media`) → `output/photos/<animal-slug>/`, mapped to the
|
||
animal by drawing anchor position.
|
||
2. **Dedup + review (stage 2)**
|
||
- Merge animals on `normalise(call-name) + DOB`, with the **Zucht as
|
||
discriminator** (Julian's ruling: Wurfchronik `[brackets]` ≡ Stammbaum
|
||
`of/von <line>` suffix — both are the breeding line; same name+DOB but
|
||
different Zucht stays two animals).
|
||
- Match animals onto Wurfchronik litters (`litterRef`) via DOB + (Vater,
|
||
Mutter) — the Pam-validated build order (chronicle litters are canonical).
|
||
- Emit a German-language `output/review-report.md` for the breeder to verify
|
||
(merges, **conflicts**, ambiguous/incomplete entries, unmapped genotype
|
||
tokens, litter data-quality warnings).
|
||
- **Nothing is loaded into the database** — stage 3 (API load) is separate and
|
||
waits on DATA-2 + FEAT-1b phase 2.
|
||
|
||
### Wurfchronik column semantics (Julian, authoritative)
|
||
|
||
`A` Wurfbezeichnung · `B` Geburtsdatum · `C` Mutter · `D` Vater (`[…]` = Zucht,
|
||
`&` = multiple sires) · `E` **survivedToGoHome** (Tabelle1 only, unlabeled —
|
||
detected positionally) · `F` Wurfstärke → `totalBorn` · `G` breakdown
|
||
`Männchen,Weibchen,TG,s` → `males/females/stillborn/diedLater` (`s` = died
|
||
after birth, before Abgabe) · last column → `note`. Validation: `E` should
|
||
equal `F − TG − s`; mismatches become German warnings in the review report
|
||
(data-quality signal, not an import blocker). A few Tabelle2 rows shift these
|
||
columns — they are read value-adaptively and flagged with a warning.
|
||
|
||
Genotypes are mapped to the frozen 8-locus contract (A C D E G P Sp Re) while
|
||
preserving everything: `genotype.mapped8locus`, `genotype.rawGenotype` (verbatim),
|
||
`genotype.unmappedTokens` (e.g. the `Uw` locus, markers `WFNZ/WP/DP`). A `-`
|
||
(unknown second allele) maps to `?`.
|
||
|
||
## Run
|
||
|
||
```sh
|
||
cd tools/import
|
||
python extract.py # uses the default source paths
|
||
python extract.py --stammbaeume "<dir>" --wurfchronik "<file.xlsx>"
|
||
```
|
||
|
||
Requires Python 3. **Re-runnable / idempotent** — re-run when more files arrive
|
||
(Wurfchronik `Teil2+`, or new charts).
|
||
|
||
## Output (`tools/import/output/`, git-ignored except the report)
|
||
|
||
| File | Contents |
|
||
|---|---|
|
||
| `animals.json` | deduped animals with genotype, parentRefs, photos, sourceFiles |
|
||
| `litters.json` | litters from the Wurfchronik |
|
||
| `photos/<slug>/…` | extracted, anchor-mapped images |
|
||
| `review-report.md` | **human review deliverable** (committed) |
|
||
|
||
## Files
|
||
|
||
- `xlsx_util.py` — dependency-free `.xlsx` reader (zip + XML): shared strings,
|
||
cells by reference, image/drawing anchors.
|
||
- `genotype.py` — genotype notation parser → 8-locus mapping + raw + unmapped.
|
||
- `extract.py` — the pipeline (stages 1–2).
|