Neuer Parser extract_contracts.py liest die ~1,4k Abgabevertrags-DOCX (\truenas\…\Verträge): er extrahiert aus dem Dokument-Body (zuverlässiger als die Dateinamen) Käufer, Tier(e), Farbschlag, Abgabedatum und Preis — robust gegen Word-Run-Splits (z. B. „F r au"/„3 0,00"); überspringt Vorlage, Abstammungsnachweise und als .docx getarnte .doc. enrich_from_contracts() in merge_and_resolve.py: Käufer werden als Kontakte (IsReceiver) angelegt/zusammengeführt; Tiere werden KONSERVATIV per Rufname (+ DOB-Jahr bei Mehrdeutigkeit) auf eigene Bestandstiere gematcht und erhalten ReceiverContactId, GoHomeDate und Status „abgegeben" — nur wo nicht bereits gesetzt; Konflikte werden geloggt, nicht überschrieben. Jede Übernahme bekommt eine Herkunfts-Zeile („Abgabe an … aus Vertrag … übernommen."). Ergebnis: 1095 Verträge → 783 Tier-Treffer (400 mehrdeutige übersprungen), 274 neue Abnehmer-Kontakte, 153 Tiere mit Abnehmer, 49 mit Abgabedatum, 23 neu „abgegeben". Keine Backend-/Frontend-Änderung nötig (Akte zeigt Abnehmer/ Abgabedatum/Herkunft bereits). SaleContract-Records bewusst nicht erzeugt (bräuchte Migration + ingest-sichere Id — späterer Schritt). Tests: test_extract_contracts.py (Dateiname/Body/Run-Split/Skip-Regeln) + alle bestehenden grün; dotnet 212. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
92 lines
4.7 KiB
Markdown
92 lines
4.7 KiB
Markdown
# GerbilManager import tooling (FEAT-8b)
|
||
|
||
One-off **migration tooling** (Python, no third-party deps) that turns Julian's
|
||
wife's hand-built spreadsheets into normalised JSON for review and, later, import.
|
||
This is *not* product code — it lives outside the app and is run manually.
|
||
|
||
See the format analysis in `FEAT-8a-format-spec.md` (Pam's hive workspace).
|
||
|
||
## What it does
|
||
|
||
`extract.py` runs **stages 1–2** of the pipeline:
|
||
|
||
1. **Extract (stage 1)**
|
||
- 10 *Stammbaum* pedigree charts → animals (name, DOB, death, Farbschlag,
|
||
genotype, breeder, positionally-reconstructed parent links, photos).
|
||
- *Wurfchronik* litter chronicle → litters (date, dam, sire, Wurfstärke,
|
||
sex breakdown, Zuchtnummer, notes). Columns are read **by header row** because
|
||
the two sheets use different schemas.
|
||
- Embedded photos (`xl/media`) → `output/photos/<animal-slug>/`, mapped to the
|
||
animal by drawing anchor position.
|
||
2. **Dedup + review (stage 2)**
|
||
- Merge animals on `normalise(call-name) + DOB`, with the **Zucht as
|
||
discriminator** (Julian's ruling: Wurfchronik `[brackets]` ≡ Stammbaum
|
||
`of/von <line>` suffix — both are the breeding line; same name+DOB but
|
||
different Zucht stays two animals).
|
||
- Match animals onto Wurfchronik litters (`litterRef`) via DOB + (Vater,
|
||
Mutter) — the Pam-validated build order (chronicle litters are canonical).
|
||
- Emit a German-language `output/review-report.md` for the breeder to verify
|
||
(merges, **conflicts**, ambiguous/incomplete entries, unmapped genotype
|
||
tokens, litter data-quality warnings).
|
||
- **Nothing is loaded into the database** — stage 3 (API load) is separate and
|
||
waits on DATA-2 + FEAT-1b phase 2.
|
||
|
||
### Wurfchronik column semantics (Julian, authoritative)
|
||
|
||
`A` Wurfbezeichnung · `B` Geburtsdatum · `C` Mutter · `D` Vater (`[…]` = Zucht,
|
||
`&` = multiple sires) · `E` **survivedToGoHome** (Tabelle1 only, unlabeled —
|
||
detected positionally) · `F` Wurfstärke → `totalBorn` · `G` breakdown
|
||
`Männchen,Weibchen,TG,s` → `males/females/stillborn/diedLater` (`s` = died
|
||
after birth, before Abgabe) · last column → `note`. Validation: `E` should
|
||
equal `F − TG − s`; mismatches become German warnings in the review report
|
||
(data-quality signal, not an import blocker). A few Tabelle2 rows shift these
|
||
columns — they are read value-adaptively and flagged with a warning.
|
||
|
||
Genotypes are mapped to the frozen 8-locus contract (A C D E G P Sp Re) while
|
||
preserving everything: `genotype.mapped8locus`, `genotype.rawGenotype` (verbatim),
|
||
`genotype.unmappedTokens` (e.g. the `Uw` locus, markers `WFNZ/WP/DP`). A `-`
|
||
(unknown second allele) maps to `?`.
|
||
|
||
## Run
|
||
|
||
```sh
|
||
cd tools/import
|
||
python extract.py # xlsx → animals.json / litters.json
|
||
python extract.py --stammbaeume "<dir>" --wurfchronik "<file.xlsx>"
|
||
python extract_docx.py # Wurfchronik-Detail.docx → docx_*.json
|
||
python extract_contracts.py # Abgabeverträge (.docx) → contracts.json
|
||
python merge_and_resolve.py # → resolved_import.json (DB-ready)
|
||
```
|
||
|
||
Requires Python 3 (zero third-party deps). **Re-runnable / idempotent** — re-run
|
||
when more files arrive (Wurfchronik `Teil2+`, new charts, or new contracts).
|
||
|
||
`extract_contracts.py` scans the breeder's sale-contract share
|
||
(`\\truenas\…\Verträge`, ~1.4k `.docx`) and emits one record per contract
|
||
(buyer, animal call-names, Farbschlag, dates, price, source filename). It skips
|
||
the blank template, `Abstammungsnachweis`/`Geburtsurkunde` documents, and any
|
||
file that is not a readable `.docx`. `merge_and_resolve.py` then conservatively
|
||
folds contracts into the resolved data: buyers become receiver `Contacts`, and
|
||
unambiguously matched gerbils get `ReceiverContactId` / `GoHomeDate` /
|
||
`Status=GivenAway` (only where not already set), with a provenance history line.
|
||
Ambiguous / unmatched animals are counted and skipped, never guessed.
|
||
|
||
## Output (`tools/import/output/`, git-ignored except the report)
|
||
|
||
| File | Contents |
|
||
|---|---|
|
||
| `animals.json` | deduped animals with genotype, parentRefs, photos, sourceFiles |
|
||
| `litters.json` | litters from the Wurfchronik |
|
||
| `docx_animals.json` / `docx_litters.json` | Wurfchronik-Detail.docx rows |
|
||
| `contracts.json` | one record per Abgabevertrag (buyer, animals, dates, price) |
|
||
| `resolved_import.json` | merged DB-ready payload consumed by `IngestResolvedService` |
|
||
| `photos/<slug>/…` | extracted, anchor-mapped images |
|
||
| `review-report.md` | **human review deliverable** (committed) |
|
||
|
||
## Files
|
||
|
||
- `xlsx_util.py` — dependency-free `.xlsx` reader (zip + XML): shared strings,
|
||
cells by reference, image/drawing anchors.
|
||
- `genotype.py` — genotype notation parser → 8-locus mapping + raw + unmapped.
|
||
- `extract.py` — the pipeline (stages 1–2).
|