Files
GerbilManager/tools/import/README.md
Gulum 5e14124322
All checks were successful
CI / Backend Tests (.NET) (push) Successful in 1m1s
CI / Frontend Tests (Node/Vite) (push) Successful in 9m35s
CI / Docker Build & Push (push) Successful in 1m18s
feat(import): Abgabeverträge (DOCX) auswerten und Tiere/Kontakte anreichern
Neuer Parser extract_contracts.py liest die ~1,4k Abgabevertrags-DOCX
(\truenas\…\Verträge): er extrahiert aus dem Dokument-Body (zuverlässiger als
die Dateinamen) Käufer, Tier(e), Farbschlag, Abgabedatum und Preis — robust
gegen Word-Run-Splits (z. B. „F r au"/„3 0,00"); überspringt Vorlage,
Abstammungsnachweise und als .docx getarnte .doc.

enrich_from_contracts() in merge_and_resolve.py: Käufer werden als Kontakte
(IsReceiver) angelegt/zusammengeführt; Tiere werden KONSERVATIV per Rufname
(+ DOB-Jahr bei Mehrdeutigkeit) auf eigene Bestandstiere gematcht und erhalten
ReceiverContactId, GoHomeDate und Status „abgegeben" — nur wo nicht bereits
gesetzt; Konflikte werden geloggt, nicht überschrieben. Jede Übernahme bekommt
eine Herkunfts-Zeile („Abgabe an … aus Vertrag … übernommen.").

Ergebnis: 1095 Verträge → 783 Tier-Treffer (400 mehrdeutige übersprungen),
274 neue Abnehmer-Kontakte, 153 Tiere mit Abnehmer, 49 mit Abgabedatum,
23 neu „abgegeben". Keine Backend-/Frontend-Änderung nötig (Akte zeigt Abnehmer/
Abgabedatum/Herkunft bereits). SaleContract-Records bewusst nicht erzeugt
(bräuchte Migration + ingest-sichere Id — späterer Schritt).

Tests: test_extract_contracts.py (Dateiname/Body/Run-Split/Skip-Regeln) + alle
bestehenden grün; dotnet 212.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 16:53:47 +02:00

92 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GerbilManager import tooling (FEAT-8b)
One-off **migration tooling** (Python, no third-party deps) that turns Julian's
wife's hand-built spreadsheets into normalised JSON for review and, later, import.
This is *not* product code — it lives outside the app and is run manually.
See the format analysis in `FEAT-8a-format-spec.md` (Pam's hive workspace).
## What it does
`extract.py` runs **stages 12** of the pipeline:
1. **Extract (stage 1)**
- 10 *Stammbaum* pedigree charts → animals (name, DOB, death, Farbschlag,
genotype, breeder, positionally-reconstructed parent links, photos).
- *Wurfchronik* litter chronicle → litters (date, dam, sire, Wurfstärke,
sex breakdown, Zuchtnummer, notes). Columns are read **by header row** because
the two sheets use different schemas.
- Embedded photos (`xl/media`) → `output/photos/<animal-slug>/`, mapped to the
animal by drawing anchor position.
2. **Dedup + review (stage 2)**
- Merge animals on `normalise(call-name) + DOB`, with the **Zucht as
discriminator** (Julian's ruling: Wurfchronik `[brackets]` ≡ Stammbaum
`of/von <line>` suffix — both are the breeding line; same name+DOB but
different Zucht stays two animals).
- Match animals onto Wurfchronik litters (`litterRef`) via DOB + (Vater,
Mutter) — the Pam-validated build order (chronicle litters are canonical).
- Emit a German-language `output/review-report.md` for the breeder to verify
(merges, **conflicts**, ambiguous/incomplete entries, unmapped genotype
tokens, litter data-quality warnings).
- **Nothing is loaded into the database** — stage 3 (API load) is separate and
waits on DATA-2 + FEAT-1b phase 2.
### Wurfchronik column semantics (Julian, authoritative)
`A` Wurfbezeichnung · `B` Geburtsdatum · `C` Mutter · `D` Vater (`[…]` = Zucht,
`&` = multiple sires) · `E` **survivedToGoHome** (Tabelle1 only, unlabeled —
detected positionally) · `F` Wurfstärke → `totalBorn` · `G` breakdown
`Männchen,Weibchen,TG,s``males/females/stillborn/diedLater` (`s` = died
after birth, before Abgabe) · last column → `note`. Validation: `E` should
equal `F TG s`; mismatches become German warnings in the review report
(data-quality signal, not an import blocker). A few Tabelle2 rows shift these
columns — they are read value-adaptively and flagged with a warning.
Genotypes are mapped to the frozen 8-locus contract (A C D E G P Sp Re) while
preserving everything: `genotype.mapped8locus`, `genotype.rawGenotype` (verbatim),
`genotype.unmappedTokens` (e.g. the `Uw` locus, markers `WFNZ/WP/DP`). A `-`
(unknown second allele) maps to `?`.
## Run
```sh
cd tools/import
python extract.py # xlsx → animals.json / litters.json
python extract.py --stammbaeume "<dir>" --wurfchronik "<file.xlsx>"
python extract_docx.py # Wurfchronik-Detail.docx → docx_*.json
python extract_contracts.py # Abgabeverträge (.docx) → contracts.json
python merge_and_resolve.py # → resolved_import.json (DB-ready)
```
Requires Python 3 (zero third-party deps). **Re-runnable / idempotent** — re-run
when more files arrive (Wurfchronik `Teil2+`, new charts, or new contracts).
`extract_contracts.py` scans the breeder's sale-contract share
(`\\truenas\…\Verträge`, ~1.4k `.docx`) and emits one record per contract
(buyer, animal call-names, Farbschlag, dates, price, source filename). It skips
the blank template, `Abstammungsnachweis`/`Geburtsurkunde` documents, and any
file that is not a readable `.docx`. `merge_and_resolve.py` then conservatively
folds contracts into the resolved data: buyers become receiver `Contacts`, and
unambiguously matched gerbils get `ReceiverContactId` / `GoHomeDate` /
`Status=GivenAway` (only where not already set), with a provenance history line.
Ambiguous / unmatched animals are counted and skipped, never guessed.
## Output (`tools/import/output/`, git-ignored except the report)
| File | Contents |
|---|---|
| `animals.json` | deduped animals with genotype, parentRefs, photos, sourceFiles |
| `litters.json` | litters from the Wurfchronik |
| `docx_animals.json` / `docx_litters.json` | Wurfchronik-Detail.docx rows |
| `contracts.json` | one record per Abgabevertrag (buyer, animals, dates, price) |
| `resolved_import.json` | merged DB-ready payload consumed by `IngestResolvedService` |
| `photos/<slug>/…` | extracted, anchor-mapped images |
| `review-report.md` | **human review deliverable** (committed) |
## Files
- `xlsx_util.py` — dependency-free `.xlsx` reader (zip + XML): shared strings,
cells by reference, image/drawing anchors.
- `genotype.py` — genotype notation parser → 8-locus mapping + raw + unmapped.
- `extract.py` — the pipeline (stages 12).