Computational Folklore Studies

FolkloreVector

Can computational representations reveal reproducible patterns of cultural continuity, mutation, diffusion, and revival in a large folklore collection — while accounting for changes in collection practices and archival composition over time?

Studied through the Journal of American Folklore, 1888–1930 — every claim below is scoped to what this specific archive can support evidence for, not to “American folklore” or “oral tradition” in general.

43
Volumes
1888–1930
Years covered
6,574
Raw OCR segments
5,490
Segments analyzed
20
NMF topics

What the analysis actually shows so far

Reported as measured, including the negative and partial results — this project’s methodology treats a null finding as real evidence, not a failure to hide. Full reasoning and result files for each claim are tracked in the project’s claims registry.

Supported

Topic structure is real, not just term-frequency noise

NMF topics on the real corpus are more stable across random seeds and document subsamples (mean Jaccard 0.678) than topics fit on a randomized-control corpus with the same vocabulary and document lengths (0.497). The subsample check carries almost all of this signal — the seed check alone barely separates real from random.

Supported

Recurring concepts show more contextual drift than stable anchors

Four pre-registered terms expected to shift in usage across 1888–1930 (ghost, war, luck, doctor) show higher decade-to-decade embedding drift than four anchor terms expected to stay stable (mother, river, moon, house) — 0.079 vs. 0.042 mean drift, with every target term exceeding every anchor term.

Negative finding

No topic shows a statistically defensible abrupt shift

After correcting for testing 20 topics simultaneously, zero topics show a significant change point in per-year prevalence — in either the raw series or a series reweighted to a common archive composition. Several topics show 1–5 change points before correction; none survive it. Whatever change is happening here looks gradual, not punctuated — or the annual bins are underpowered to tell the difference. Both explanations are still live.

Partial / complicating

Semantic search beats lexical search — but not cleanly

On the one hand-confirmed cross-vocabulary motif match in a first small validation round, semantic search ranked the true match roughly 10x higher than lexical search. But two of three deliberately-chosen hard-negative pairs — sharing only a title or a character name, not an actual narrative — ranked just as well or better under both methods. Low retrieval rank alone is not yet good evidence of real motif detection here.

The archive explorer, semantic atlas, timeline, motif network, and full methods/paper sections are under active construction and not yet part of this deployment.