Files
GerbilManager/tools/import
Gulum 0b99cfd2bd
Some checks failed
CI / Backend Tests (.NET) (push) Successful in 1m16s
CI / Frontend Tests (Node/Vite) (push) Successful in 9m39s
CI / Docker Build & Push (push) Successful in 14m0s
CI / Deploy auf TrueNAS (Custom App) (push) Failing after 3s
feat(import): neue Stammbäume (12 Charts) + litterChildren.add + EMF-Fotos überspringen
Die Züchterin hat 24 neue/aktualisierte Stammbaum-xlsx geliefert (Ordner
"neuestammbäume"); sie liegen jetzt im kanonischen Quellverzeichnis
Sttammbäume (12 neue Charts, 7 aktualisierte, 3 identisch, das inhaltsgleiche
"Picus Son (2)" ausgelassen). Prod ist per Upload-Ingest aktualisiert:
2372 -> 2451 Tiere, 916 -> 965 Würfe, 432 -> 507 Fotos, 2198 -> 2275 Tiere
mit Geburtsdatum. Overrides/verified-Zeilen, manuelle Tiere und Tickets
haben den Ingest unverändert überlebt.

Zwei Datenfehler, die die neuen Charts aufgedeckt haben — datengetrieben und
re-ingest-stabil gefixt statt an der globalen Heuristik zu drehen:

- litterChildren kennt jetzt `add` [Name | {name, dob}] als Gegenstück zu
  `keep`: hängt ein Jungtier an DIESEN Wurf und entfernt den alten Wurf, wenn
  er dadurch kinderlos UND virtuell ist. Nötig, weil "Pukas Kids" Akanes
  Eltern komplett UNTER ihren Block setzt (N80 Roni = Vater, N81 Fumi =
  Mutter) — _reconstruct_parents griff eine Zeile zu hoch, paarte Irish
  Coffee (Bonapartes Mutter) mit Roni und riss Akane aus dem Z21-Wurf in
  einen Phantom-Wurf, der in der Wurfchronik auftauchte (Ticket 88389f8e).
- Merle: durch das neue Geburtsdatum (18.06.2023) mergt der addAnimals-Stub
  in den Chart-Datensatz und verliert dabei sein isResident -> expliziter
  resolutions-Override (Ticket 36a3fcde/a8f11ac0, Züchterin: Zuchttier).

Außerdem: Excel legt neben Fotos teils EMF/WMF-Vektorvorschauen ab, die
Browser nicht darstellen können (kaputte Bildkachel in der Tier-Akte) ->
extract._attach_photos überspringt .emf/.wmf (5 Fotos betroffen).

Regressionstests für alle drei Punkte; alle Python-Suites grün.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 22:27:18 +02:00
..

GerbilManager import tooling (FEAT-8b)

One-off migration tooling (Python, no third-party deps) that turns Julian's wife's hand-built spreadsheets into normalised JSON for review and, later, import. This is not product code — it lives outside the app and is run manually.

See the format analysis in FEAT-8a-format-spec.md (Pam's hive workspace).

What it does

extract.py runs stages 12 of the pipeline:

  1. Extract (stage 1)
    • 10 Stammbaum pedigree charts → animals (name, DOB, death, Farbschlag, genotype, breeder, positionally-reconstructed parent links, photos).
    • Wurfchronik litter chronicle → litters (date, dam, sire, Wurfstärke, sex breakdown, Zuchtnummer, notes). Columns are read by header row because the two sheets use different schemas.
    • Embedded photos (xl/media) → output/photos/<animal-slug>/, mapped to the animal by drawing anchor position.
  2. Dedup + review (stage 2)
    • Merge animals on normalise(call-name) + DOB, with the Zucht as discriminator (Julian's ruling: Wurfchronik [brackets] ≡ Stammbaum of/von <line> suffix — both are the breeding line; same name+DOB but different Zucht stays two animals).
    • Match animals onto Wurfchronik litters (litterRef) via DOB + (Vater, Mutter) — the Pam-validated build order (chronicle litters are canonical).
    • Emit a German-language output/review-report.md for the breeder to verify (merges, conflicts, ambiguous/incomplete entries, unmapped genotype tokens, litter data-quality warnings).
    • Nothing is loaded into the database — stage 3 (API load) is separate and waits on DATA-2 + FEAT-1b phase 2.

Wurfchronik column semantics (Julian, authoritative)

A Wurfbezeichnung · B Geburtsdatum · C Mutter · D Vater ([…] = Zucht, & = multiple sires) · E survivedToGoHome (Tabelle1 only, unlabeled — detected positionally) · F Wurfstärke → totalBorn · G breakdown Männchen,Weibchen,TG,smales/females/stillborn/diedLater (s = died after birth, before Abgabe) · last column → note. Validation: E should equal F TG s; mismatches become German warnings in the review report (data-quality signal, not an import blocker). A few Tabelle2 rows shift these columns — they are read value-adaptively and flagged with a warning.

Genotypes are mapped to the frozen 8-locus contract (A C D E G P Sp Re) while preserving everything: genotype.mapped8locus, genotype.rawGenotype (verbatim), genotype.unmappedTokens (e.g. the Uw locus, markers WFNZ/WP/DP). A - (unknown second allele) maps to ?.

Run

cd tools/import
python extract.py                       # xlsx → animals.json / litters.json
python extract.py --stammbaeume "<dir>" --wurfchronik "<file.xlsx>"
python extract_docx.py                  # Wurfchronik-Detail.docx → docx_*.json
python extract_contracts.py             # Abgabeverträge (.docx) → contracts.json
python merge_and_resolve.py             # → resolved_import.json (DB-ready)

Requires Python 3 (zero third-party deps). Re-runnable / idempotent — re-run when more files arrive (Wurfchronik Teil2+, new charts, or new contracts).

extract_contracts.py scans the breeder's sale-contract share (\\truenas\…\Verträge, ~1.4k .docx) and emits one record per contract (buyer, animal call-names, Farbschlag, dates, price, source filename). It skips the blank template, Abstammungsnachweis/Geburtsurkunde documents, and any file that is not a readable .docx. merge_and_resolve.py then conservatively folds contracts into the resolved data: buyers become receiver Contacts, and unambiguously matched gerbils get ReceiverContactId / GoHomeDate / Status=GivenAway (only where not already set), with a provenance history line. Ambiguous / unmatched animals are counted and skipped, never guessed.

Output (tools/import/output/, git-ignored except the report)

File Contents
animals.json deduped animals with genotype, parentRefs, photos, sourceFiles
litters.json litters from the Wurfchronik
docx_animals.json / docx_litters.json Wurfchronik-Detail.docx rows
contracts.json one record per Abgabevertrag (buyer, animals, dates, price)
resolved_import.json merged DB-ready payload consumed by IngestResolvedService
photos/<slug>/… extracted, anchor-mapped images
review-report.md human review deliverable (committed)

Files

  • xlsx_util.py — dependency-free .xlsx reader (zip + XML): shared strings, cells by reference, image/drawing anchors.
  • genotype.py — genotype notation parser → 8-locus mapping + raw + unmapped.
  • extract.py — the pipeline (stages 12).