# Retained content-topic measurement

The recoverable sample does not establish a dominant original subject for any of the 16 atlas portraits. It does establish a leading **workflow in the selected DSE sample**: task timing and coordination occurs in 10 of 16 distinct retained latest-revision artifact bodies (62.5%). That is a description of this recovery sample, not a community-wide topic estimate.

The Senior Data Scientist and Senior Data Engineer principles applied here are DS-1/2/4/8 and DE-1/3/6/7/8. This is the measurement contribution to the parent persona review, not a separately constituted or certified full panel.

## Instrument and denominator

Run `python3 reproduce.py` from this directory. Python standard library only; all sources are read offline. The script verifies every annotated raw body against both the original manifest hash and the annotation hash; any discrepancy fails. It selects the highest **recovered** numeric revision per exact namespace and artifact, then removes exact duplicate raw-body hashes within each namespace. No revisions from different namespaces are merged. It recomputes `result.json` deterministically from `annotations.csv` and retained inputs. No raw source body, credential or executable markup is copied here.

There are 37 source targets, 18 verified revision bodies, 16 distinct artifacts after revision selection, zero exact duplicate bodies after that selection, and 19 unrecovered or unverifiable targets. Every eligible body is from `wikiservice.at/dse`. Different revisions are not independent conversations; neither artifact count nor tracker observation count is an agent or population count.

Annotations use conservative manual reading of sanitized lexical evidence and exact artifact context. Original subject and coordination workflow are separate columns. A date or subject in a page name alone cannot establish the retained body's original subject. Raw instructions were never executed. This is single-coder exploratory annotation, not validated semantic classification; there is no inter-rater reliability estimate. The code replays the assigned labels; it does not prove their semantic correctness.

## Observed sample

| Original subject | Bodies / 16 |
|---|---:|
| Unknown original subject | 14 |
| Agriculture and crop data | 1 |
| Digital archives and library objects | 1 |

| Coordination workflow | Bodies / 16 |
|---|---:|
| Task timing and coordination | 10 |
| Resource pointers and retrieval | 5 |
| Markup experiment | 1 |

Independent review rejected the CVD subject label: the only cue is an identifier substring, with no standalone health term. It is now unknown. The earlier grocery revision explicitly concerns workforce/grocery data, but the later retained revision is a short coordination relay. Counting both revisions or transferring the earlier subject to the later body changes the apparent mix. We preserve that earlier annotation for audit but exclude it from the latest-revision denominator.

No confidence interval is supplied for community prevalence because these are targeted recoveries selected for coordination and safety mechanisms, with unknown inclusion probabilities. A binomial interval would hide the selection bias. No community-wide share is estimated. A null share means unknown, not zero.

## Coverage and evidence levels

`result.json` contains all 16 exact `data.js` ids and the requested `measured_sample_n`, `dominant_topic`, `dominant_n`, `share`, `status`, `caveat` and `source_refs` fields. DSE has a measured denominator of 16; the other 15 have zero eligible original bodies in this pinned manifest. Their reported themes remain secondary evidence and are not converted into topic frequencies. This does not mean those venues have no original artifacts elsewhere or no activity.

The inventory contains 154 venue rows and 29 system leads, not 183 independent civilizations. Counts and all unmeasured venue ids/system names are reproduced in the JSON coverage section. Parent hosts, aliases, data targets, archives, candidates and designed systems remain their original categories. Tarcseh and k4be remain a display grouping only; no shared denominator is inferred. Etherpad pads retain separate identities; RubyGems metadata is not code execution; Xinzhai ciphertext yields no plaintext topic inference. Separate Artifactory and AISI incident reports do not become public-community samples.

The 146 repository artifact records are event summaries, including 120 Iowa-family records and 26 references. They are excluded from original-body prevalence. The independent reviewer additionally re-derived 120 Iowa metadata objects and 110 source-assigned body groups. Neither number denotes original bodies independently hashed here, independent conversations, authenticated agents, nor general Linuxiarz topic prevalence. The Iowa focus was built into the selected family.

## Time boundary

The input is retained research from the September 7 local session; the manifest captures crossed UTC midnight into September 8. Exact capture endpoints are in `result.json`. This is not a sample of messages written yesterday. Source event time is unknown for this measured set; page names containing dates cannot repair it. The inventory tracker first/last timestamps are metadata coverage, not body creation timestamps. No live acquisition or current community-state claim was made.

## Findings for parent synthesis

[P1] DS-1, DS-2, DS-8: Targeted body recovery cannot identify a community's dominant original subject.
Evidence: `python3 reproduce.py` produces 18 verified revisions, 16 bodies, 14 unknown original subjects. Grade (a) for recomputed counts; semantic coding remains provisional.
Impact: Assigning a specialist economy as a measured fact would overstate the evidence.
Fix: The integration fields leave dominance null and expose the separately scoped workflow measurement.

[P1] DS-4, DE-8: Fifteen portraits have no eligible original-body sample in this manifest.
Evidence: every verified revision has the `dse~` namespace; all 16 portrait records are present in `result.json`. Grade (a) for coverage; secondary themes are grade (b).
Impact: Investigator prose could otherwise masquerade as measured community discourse.
Fix: Render secondary themes and unknown dominance visibly; acquisition outside the pinned manifest remains a future sampling task.

[P2] DE-1, DE-6, DE-7: Reproducibility is hash-bound and deterministic, while annotation reliability is unmeasured.
Evidence: `reproduce.py` verifies raw byte hashes and explicit revision identities before counting. Grade (a) for replay mechanism, pending independent semantic review.
Impact: Replay detects source drift but cannot validate a mislabeled topic.
Fix: Keep annotations inspectable and distinguish replay verification from independent coding validation.

## Broader retained-corpus check

The broader crawl manifest lists 462 files. The corpus manifest explicitly labels all 420 artifacts as investigator reports. The remaining crawl entries are report indexes, provenance/claims/relations, Iowa object metadata, venue/incident registries and investigator HTML pages. No whole original-body corpus export corresponding to the reported 14,591 revisions appears in either manifest or the inspected retained data/cache filename inventory. The 18 exact-revision recovery bodies remain the bounded eligible set used here; this is an inventory search result, not proof that no additional source body exists anywhere on disk. The larger publisher corpus count is not substituted for locally measured bodies.

## Recovered-history sensitivity

The artifact-level union of explicit subject evidence across all 18 recovered revisions still has exactly 16 distinct artifacts as denominator. It yields 3 artifacts with a known historical subject and 13 without one: workforce/grocery 1, agriculture/crops 1, digital archives 1. Latest-revision coding yields 2 known and 14 unknown. The sole difference is `dse~DataUSAGrocerySequenceCollab2027`: revision 1 supplies the explicit workforce/grocery subject, while revision 20 supplies generic coordination. Thus unknown in the latest body does not mean no subject was established in retained history.

History labels are unioned within each exact namespace/artifact. Multi-label counts are allowed and may sum above the denominator in future inputs, although no artifact has multiple known subjects in this set. This does not count revisions as independent observations and does not claim access to complete revision history. The machine-readable comparison is `history_subject_sensitivity` in `result.json`; both policies leave community dominance null.
