Second and third rounds on the dossier pipeline, reviewed locally before commit. Retrieval - wayback_pass.py: retries every BLOCKED/DEAD/THIN/ERROR linked source through the Wayback Machine; blocks carry the snapshot date, archive URL and the live verdict. 143 slots retried: 93 fetched, 37 no snapshot, 12 script shells, 1 PDF. - fetch_sources.py: fetch_wayback() helper. - add_extra_sources.py: files staged extra_* finds (web search, not on the DB record) into the store with a provenance note. - Cases re-sourced by web search: 248 (Oslo: Document.no, Avisa Oslo, NRK; the linked Le Parisien piece is case 242 and is marked OFF-CASE), 258 Kharkiv, 260 Singapore, 358 Bangkok, 365 Las Vegas, 388 Verneuil-sur-Seine, 390 Zoersel, 430 Homestead. 341 of 364 cases now hold a fetched article; 11, 66 and 399 have no public text source (podcast, police video, direct victim report). Trimming - auto_trim.py: case-anchored furniture cut for UNTRIMMED blocks. Finds the body run that mentions the case, merges across subheadings and short furniture gaps, drops teasers, share bars, date/URL/caption lines and subscription pitches; refuses pages with no record term or almost no body. 302 blocks trimmed; 16 left on auto-trim-review.md (7 OFF-CASE suspects). Every cut is labelled AUTO-TRIMMED, UNREVIEWED in the dossier; the full extract stays in staging/. Decisions with anchors in trim-decisions-auto.json. - build_cases.py / verify_all.py: banners for auto-trimmed and archived blocks, staged-file check extended to Wayback blocks, flag legend under the record. Detention rule - README "Definitions": the DB field `kidnappings` is the DETENTION violence type (victim, guard, staff or relative held to force submission or execute the theft), distinct from the Kidnapping scenario (taken away and held). - detention-flag-review.md / detention-flag-corrections.json: 40 records reviewed with evidence; 30 set-to-1 proposals accepted by the owner on 2026-09-06 (listed in corrections-approved.md), 10 still open. - apply_detention_wording.py: "Violence Used" in the 70 reviewed summaries now names detention explicitly (62 of 70 labelled), supported by the summary's own text; idempotent; supersedes the batch scripts' wording. Worklist and docs - sofia-worklist.md: 248 decisions, detention rule item replacing the old "no abduction" item, fresh-search section for the textless cases. - README-START-HERE.md: progress notes, run commands, next steps. - Bug fixed in passing: Wayback blocks stored in-memory text with carriage returns; now stored as read back from disk. verify_all.py PASSES (2321 checks). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GZZENdTLzNGsbNy4DyF1yt
272 lines
17 KiB
Markdown
272 lines
17 KiB
Markdown
# Gart K&R Database - Handoff (read this first)
|
|
|
|
Written 2026-09-02. This folder is a self-contained bundle. Everything the next AI needs to
|
|
continue is inside it, and all script paths are relative, so the folder can be moved or zipped
|
|
and still run.
|
|
|
|
Read order: (1) this file, (2) merge-output/sofia-worklist.md (the human's manual task list).
|
|
|
|
--------------------------------------------------------------------------------
|
|
## Folder layout
|
|
|
|
Gart-KR-handoff/
|
|
README-START-HERE.md <- this file
|
|
source-checker-agent.md <- the read-only reviewer agent (install in .claude/agents/)
|
|
attacks-export-Gart-website.json <- SOURCE: the live DB export (read-only)
|
|
KR-reports_analysis_Sofi.xlsx <- SOURCE: the old article spreadsheet (read-only)
|
|
merge-output/
|
|
sources/ <- THE DELIVERABLE: 364 .txt dossiers (build_cases writes here)
|
|
build_cases.py <- rebuilds all 364 .txt files
|
|
verify_all.py <- deterministic audit (must say PASS)
|
|
survey_summaries.py <- lists which cases still have short summaries
|
|
build_overrides_batch1..7.py <- the summary rewrites, one script per batch
|
|
dump_cell.py <- helper: prints a spreadsheet cell line-by-line
|
|
summary-overrides.json <- the improved summaries (build input)
|
|
sheet-trim-decisions.json <- article furniture cuts (build input)
|
|
sourced-articles.json <- fetched-article store (build input)
|
|
sofia-worklist.md <- the human's manual to-do list
|
|
(other .py and .json files are finished-phase artifacts, not needed to continue)
|
|
|
|
Scripts live in merge-output/ and must be run from there (they find the two source files in
|
|
the parent folder automatically). Always call them with the full path, for example:
|
|
python "<path-to>/Gart-KR-handoff/merge-output/build_cases.py" --all
|
|
|
|
Requirements to run this:
|
|
- Python 3 with the openpyxl package (pip install openpyxl) - build_cases.py reads the .xlsx.
|
|
- The review step uses a read-only fact-checking sub-agent. Its definition is included here as
|
|
source-checker-agent.md (bundle root). In Claude Code, drop it into your project's
|
|
.claude/agents/ folder so you can invoke it; otherwise any capable read-only checker works -
|
|
just feed it the source-checker prompt (further down this doc) plus the batch's file list.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## What this project is
|
|
|
|
Gart's K&R attack database lists 364 real kidnap/robbery/extortion incidents against
|
|
cryptocurrency holders. For each case we merged the structured DB record with any saved news
|
|
article, and produce one readable .txt dossier per case. There is no app, no build system, no
|
|
tests - just Python scripts that regenerate the .txt files, plus Markdown notes.
|
|
|
|
### Hard rules (do not break these)
|
|
1. Never edit the live database. attacks-export-Gart-website.json and the .xlsx are READ-ONLY.
|
|
Any DB correction goes to sofia-worklist.md for Sofia to apply herself.
|
|
2. Never fabricate. Every statement must trace to that case's own DB record or its article text.
|
|
If a source does not say it, do not write it. Record-only cases stay short - do not pad them.
|
|
3. Do not translate articles into the archive. Non-English article text is stored verbatim.
|
|
You MAY read a foreign-language article to write an English summary (that is derived facts,
|
|
not a stored translation), but never write a translation into a stored file.
|
|
4. Work in descending ID order (nothing gets skipped).
|
|
5. Every record-vs-article contradiction is a note for Sofia - log it, do not "resolve" it by
|
|
guessing which side is right.
|
|
|
|
### Environment
|
|
- Windows / PowerShell, with a Bash tool available (Git Bash).
|
|
- In the Bash tool, a "cd" inside a compound command triggers a permission prompt - always use
|
|
full paths instead.
|
|
- Any Python script that prints non-ASCII must wrap stdout, or it crashes on Windows:
|
|
import sys, io
|
|
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8", errors="replace")
|
|
Put that inside main() if another script imports the file.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## Definitions: detention (violence type) vs Kidnapping (scenario)
|
|
|
|
Added 2026-09-06 after the flag turned out to be mislabelled. The rule for every file,
|
|
summary and correction from now on:
|
|
|
|
**Detention** is a VIOLENCE TYPE. The DB flag `kidnappings` (sheet column "Kidnappings") is
|
|
the detention flag and should be renamed `detention` on the DB side. It is 1 when the
|
|
attackers had physical control of the victim at any point, in any place, for any length of
|
|
time: tied, taped, handcuffed or zip-tied; locked in a room or vehicle; held at gunpoint
|
|
while forced to unlock or transfer; taken away in a car; held under a fake arrest; kept for
|
|
hours or days. It is 0 when control was never gained: a snatch or mugging in passing, a
|
|
break-in of empty premises, threats or extortion at a distance, a plot foiled before contact,
|
|
an attempt the victim fled or fought off before being restrained. Judgment line: a victim
|
|
held at gunpoint for the minutes it takes to force a transfer counts as detained; a victim
|
|
punched and left within seconds does not. Who was held does not change the flag: if detention
|
|
was used on anyone to force submission or to execute the theft (guards, staff, a concierge, a
|
|
relative), the flag is 1 even when the record's named victim was absent (owner's rule, 2026-09-06).
|
|
|
|
**Kidnapping** is a SCENARIO, one of eleven attack patterns. It means the victim was taken
|
|
from where they were to a place under the attackers' control and held there, usually for a
|
|
ransom or a forced transfer. Express Kidnapping is the short, ATM-style variant. A Kidnapping
|
|
scenario always involves detention, and usually other violence types too: weapon threats,
|
|
assault, torture, drugging. Detention also occurs in other scenarios without making them
|
|
kidnappings: a family tied up in a Home Invasion, a P2P buyer handcuffed in a hotel room, a
|
|
fake arrest under Authority Misuse, a victim zip-tied in a Dodged Attack that was interrupted.
|
|
|
|
**Wording in summaries.** "Violence Used" names the violence types, and names detention
|
|
explicitly whenever it occurred ("Detention — held at gunpoint for about an hour"). The
|
|
words kidnapped / abducted are reserved for a victim being taken away; a victim held where
|
|
they were is "held", "detained", "tied up", not "kidnapped". Foiled attempts say
|
|
"No detention — the attempt was foiled" only when the source says so.
|
|
|
|
Per-case flag review against this rule: merge-output/detention-flag-review.md (proposals for
|
|
Sofia, with the evidence line for each). The DB export stays untouched.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## The deliverable, and what still needs the human
|
|
|
|
Two different kinds of change came out of this work. They are NOT the same:
|
|
|
|
1. The improved SUMMARIES are already baked into the .txt files. build_cases.py renders each
|
|
rewritten summary into its dossier under the banner "AI SUMMARY [revised on review ...]".
|
|
So if Sofia uploads the .txt files as the DB's dossier source, the summaries are upgraded.
|
|
There is no separate "apply-list" to work through for summaries. summary-overrides.json is
|
|
just the build INPUT that injects them; it is not a manual chore.
|
|
|
|
2. The DB FIELD / FLAG corrections in sofia-worklist.md are NOT in the .txt files. The
|
|
DATABASE RECORD block in each dossier is printed verbatim from the live DB, so it still shows
|
|
the old, wrong field values (for example case 353 still says "100 million rubles", case 219
|
|
still says "beaten with a stick"). Uploading the .txt does NOT fix those. They are advisory
|
|
notes for Sofia to apply on the DB side if and when she wants. She said she will do these last.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## What is already FINISHED
|
|
|
|
- Matching: all 364 records matched to their spreadsheet rows / fetched sources.
|
|
- Article merge: spreadsheet and web-fetched articles are in the dossiers.
|
|
- Article trimming: site furniture removed from fetched and spreadsheet articles; every kept
|
|
line is a verbatim line of the source; provenance hashes preserved. Reviewed twice.
|
|
- Source-coverage reconciliation: 344 of 364 cases have a usable source; 20 have none (all
|
|
listed on the worklist). No case silently lost its last source.
|
|
- verify_all.py PASSES (896 checks). Confirmed from this bundle's new location.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## The task in progress: summary enrichment
|
|
|
|
Goal: bring the short/weak bulk-import summaries up to the standard of the newest cases
|
|
(447-462), which have six labelled categories and run roughly 970-2100 characters:
|
|
Victims / Attackers / Attack Method / Violence Used / Crypto Demanded/Stolen / Status
|
|
|
|
A survey (survey_summaries.py) found 256 of 364 cases under the bar (under 400 chars or fewer
|
|
than 4 categories). These are the targets, worked in descending ID order.
|
|
|
|
### How it works
|
|
- Rewrites live in summary-overrides.json, keyed by DB id (string).
|
|
- build_cases.py loads it and renders the rewrite in place of the old bulk summary.
|
|
- The live DB export is never touched.
|
|
|
|
### Method, per case
|
|
- Read the whole dossier: the DATABASE RECORD block AND every ARTICLE / FETCHED ARTICLE section
|
|
(any language).
|
|
- Write the six categories from ONLY what the record + that case's articles support. No repeated
|
|
facts across categories. Keep the source's own hedging ("allegedly", "police said", "he says").
|
|
- Record-only or wrong-source cases (dead/blocked/thin link): restructure the record's
|
|
description + notes + flags into the six categories, faithfully, and stay short. Do not invent.
|
|
- Round-up articles (one article, several incidents): use ONLY the passage about THIS case;
|
|
drop the other incidents and any trend-commentary tail.
|
|
- Foiled / "dodged" attacks with no one harmed: "Violence Used: None - the attempt was foiled."
|
|
- Name detention explicitly in "Violence Used" whenever the victim was held (see Definitions);
|
|
use kidnapped / abducted only for a victim taken away.
|
|
|
|
### Workflow per batch (about 10 cases)
|
|
1. Write build_overrides_batchN.py (copy an existing one; it does existing.update(B) then writes
|
|
summary-overrides.json; escape non-ASCII with \\uXXXX). Run it.
|
|
2. Run build_cases.py --all, then verify_all.py (must say PASS).
|
|
3. Run the source-checker sub-agent (read-only; tell it NOT to fetch) to fact-check the batch's
|
|
summaries against each file's own record + article. This is Sofia's explicit requirement.
|
|
4. Fix every fair flag in the batch script, rebuild, verify. The batch is then locked.
|
|
5. Add any DB field/source problems you find to sofia-worklist.md.
|
|
|
|
### The source-checker prompt that works
|
|
"Read-only fact-check. Do NOT fetch. For each file, verify every claim in the revised summary
|
|
(under the 'revised on review' banner) is supported by the DATABASE RECORD block and/or the
|
|
embedded article text in the SAME file. Flag FABRICATED / CONTRADICTED / OVERSTATED (note:
|
|
source-matching hedges like 'allegedly' are correct, not flags). Tell it about round-up files,
|
|
foiled plots, and any known wrong-source case so it does not over-flag."
|
|
|
|
### The pitfalls the reviewer keeps catching (avoid these up front)
|
|
1. Dropped hedges - keep "allegedly / he believes / police said"; do not state a claim as fact.
|
|
2. Inferred negatives - do not write "no arrests reported" when the source is just silent on it.
|
|
3. Unsupported cross-links - do not tie a case to another case unless THIS file says so.
|
|
4. Added specifics not in the file - no invented brand ("Toyota" when it says only "Vigo"),
|
|
residency, or family composition ("two parents" when only "the mother" is mentioned).
|
|
5. Number / currency / date math - get units exact; when the DB date and the article date
|
|
disagree, say "late March" and log the conflict rather than silently picking one.
|
|
6. Mis-attributed timestamps - a time in the article belongs to a specific moment; do not move it.
|
|
7. Round-up / tail exclusion - use only this case's passage (the reviewer expects this).
|
|
8. Follow the article over a wrong DB field, and log the conflict on the worklist. Exception: a
|
|
record-only / wrong-source case has no article - write from the record and say so.
|
|
9. Naming - only name a person the linked source names. If the source withholds the victim's
|
|
name but the DB carries one, describe the victim generically in the summary body.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## Progress
|
|
|
|
Done and reviewer-verified: batches 1-7, 70 of 256 cases. IDs completed:
|
|
B1: 372,371,370,369,368,367,365,364,362,353
|
|
B2: 350,349,340,334,323,322,321,317,315,312
|
|
B3: 309,308,307,304,303,301,299,298,297,296
|
|
B4: 293,292,290,288,285,279,273,272,271,270
|
|
B5: 266,263,256,253,252,249,248,247,246,245
|
|
B6: 244,243,242,240,238,237,236,235,234,233
|
|
B7: 232,231,230,229,228,224,222,221,219,218
|
|
|
|
summary-overrides.json holds 70 entries. verify_all.py PASSES.
|
|
|
|
Remaining: continue descending from ID 217 (217,216,215,214,213,212,211,210,208,207 ...) down to
|
|
case 1 - about 186 short cases, roughly 19 more batches. Get the exact remaining list with:
|
|
python ".../merge-output/survey_summaries.py"
|
|
and take the ids flagged SHORT / unstructured, descending.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## 2026-09-05 update: linked URLs fetched for the 252 spreadsheet-only cases
|
|
|
|
Until this pass, fetching had only run on cases 357+ and a handful of older ones. The other
|
|
252 cases held only the text pasted into the spreadsheet, which is often just the lede.
|
|
stage_all.py now takes --include-sheet-cases, which fetches every linked URL for those cases
|
|
too. Run from merge-output/ with:
|
|
python stage_all.py --include-sheet-cases
|
|
It is resumable: anything already in sourced-articles.json is skipped.
|
|
|
|
Result of the run (all 252 cases processed):
|
|
- 179 cases gained at least one fetched article; 73 got nothing (site blocked the request,
|
|
link dead, page needs JavaScript, or the link is a video/social post).
|
|
- 296 URL slots tried: 188 FETCHED, 49 BLOCKED, 18 THIN, 9 DEAD, 4 NOT AN ARTICLE, 3 ERROR.
|
|
- Every fetched block is filed UNTRIMMED (banner in the dossier) and sits beside the sheet
|
|
paste, which is kept. Trimming is still to do, same as for cases 357+.
|
|
- verify_all.py PASSES (1836 checks).
|
|
|
|
Second pass, same day: Wayback Machine retrieval and automatic trimming.
|
|
|
|
- wayback_pass.py retries every linked source the live fetch could not read (BLOCKED,
|
|
DEAD, THIN, ERROR) through the Wayback Machine. Run from merge-output/:
|
|
python wayback_pass.py
|
|
Result over 143 failed slots: 93 FETCHED, 37 no snapshot, 12 THIN (script shell), 1 PDF.
|
|
Every archived block is labelled with the snapshot date and archive URL in the dossier
|
|
and keeps the live verdict in its note. 341 of 364 cases now hold a fetched article;
|
|
3 cases still have no article text at all: 11 (podcast only), 66 (police Facebook video only),
|
|
399 (direct victim report to Gart Research, no public source). The other seven were sourced by
|
|
web search on 2026-09-06; add_extra_sources.py files such finds into the store.
|
|
- auto_trim.py cuts site furniture from UNTRIMMED blocks: it finds the paragraph run that
|
|
mentions the case (terms from the DB record), widens it across short gaps so ledes and
|
|
closing sentences survive, and drops teasers, share bars, date/URL/caption lines and
|
|
subscription pitches. Run from merge-output/:
|
|
python auto_trim.py (add --reset to redo every automatic cut from staging/)
|
|
Result: 288 blocks in 252 cases trimmed, 1.44M -> 0.99M chars. These are MACHINE cuts:
|
|
the dossier banner says AUTO-TRIMMED, UNREVIEWED, and the full extract stays in staging/.
|
|
Decisions with anchors are in trim-decisions-auto.json.
|
|
- auto-trim-review.md lists the 16 blocks left UNTRIMMED and the 15 trimmed on a weak
|
|
anchor. The 7 OFF-CASE entries (no record term on the page) are wrong-source candidates;
|
|
case 248 is on it, which matches the known Le Parisien mismatch.
|
|
- The trimmer removes furniture only. It does NOT separate other incidents inside a
|
|
round-up article; that is still a read-through job (see the round-up bucket on the
|
|
worklist).
|
|
- Case 248: Oslo sources found by web search on 2026-09-05 and filed as extra_1..5 (see the worklist entry). verify_all.py PASSES (2321 checks). Not committed yet; review before pushing.
|
|
|
|
--------------------------------------------------------------------------------
|
|
## Quick-start checklist for the next AI
|
|
1. Read this file and merge-output/sofia-worklist.md.
|
|
2. Confirm the bundle is healthy:
|
|
python ".../merge-output/verify_all.py" (should say PASS)
|
|
summary-overrides.json should have 70 entries.
|
|
3. Pick the next ~10 short cases descending from ID 217 (use survey_summaries.py). Read each
|
|
dossier in full. Write build_overrides_batch8.py (copy batch 7's structure). Run it, then
|
|
build_cases.py --all, then verify_all.py.
|
|
4. Run the source-checker agent on the batch (read-only, no fetch). Apply fair flags; rebuild;
|
|
verify. Then the batch is locked.
|
|
5. Repeat down to case 1.
|
|
6. Keep adding DB field/source problems to sofia-worklist.md. Never edit the DB.
|
|
7. When all summaries are done, the deliverables are: the sources/*.txt (already carry the
|
|
improved summaries) and sofia-worklist.md (Sofia's manual DB fixes, done last on her side).
|