Methodology
How the lore was compiled, verified, and maintained.
Data Sources
The Whispering Lore dataset draws from a wide range of sources across folklore studies, anthropology, and comparative mythology:
- Published folklore collections and academic anthologies (e.g., Asbjørnsen & Moe, the Brothers Grimm, Y.T. Kawamura)
- Regional mythographic surveys and encyclopedias of world mythology
- Peer-reviewed journal articles in folklore studies and cultural anthropology
- Primary oral tradition transcripts where available in published form
- Indigenous cultural heritage materials published with community permission
Each entry includes a source field naming the specific reference used. Entries marked source_quality: "verified" have been cross-checked against at least two independent scholarly sources. Entries marked "unknown" are pending verification.
Inclusion Criteria
A creature or story is included if it meets at least one of:
- Documented in a published folklore or mythology source
- Part of an identifiable cultural tradition with oral or literary transmission
- Referenced in multiple independent sources (for cross-cultural entries)
Entries derived from modern fiction, video games, or film-only creations are excluded unless they have roots in traditional folklore. Contemporary legends and urban folklore are included where well-documented.
Data Structure
Each creature entry contains 25+ fields organized into categories:
- Identity: name, slug, alternative names, type classification
- Geography: country, region, cultural tradition
- Description: appearance, behavior, habitat, cultural significance
- Relations: related creatures, parent species, associated stories
- Metadata: source quality, last updated, verification status
Editorial Standards
- Entries are written in a neutral, descriptive tone that respects the source culture
- Alternative names and transliterations are preserved alongside primary names
- Regional and cultural variants are noted where documented
- Ambiguous or contested folklore is flagged rather than flattened into a single narrative
- AI-assisted text generation has been used for some summary expansions and is clearly marked in the metadata; such entries are explicitly tagged and are being progressively replaced with verified content
Verification Process
Verification is an ongoing process. Entries pass through these stages:
- Stage 1 — Initial compilation: Data collected from published sources, formatted into the standard schema
- Stage 2 — Cross-reference: Creature-to-story and story-to-creature references checked for accuracy
- Stage 3 — Source verification: Named sources checked against library catalogs and academic databases
- Stage 4 — Expert review: Subject-matter folklorists review entries for cultural accuracy
As of the current dataset version, the database contains 3,668 creatures and 2,185 stories — a combined total of 5,853 entries. All entries have source quality and source type classifications.
Current source quality distribution (creatures):
- Expert: 583 entries — reviewed by subject-matter specialists
- Verified: 60 entries — independently confirmed against primary sources
- Researched: 94 entries — traced to named academic or archival sources
- Well-documented: 161 entries — supported by multiple independent sources
- Academic: 90 entries — based on peer-reviewed academic literature
- Documented: 19 entries — recorded in named secondary sources
- Primary: 4 entries — sourced from first-hand ethnographic fieldwork
- Good: 1,720 entries — from general references
- Fair: 447 entries — from general references or secondary sources
- Poor: 490 entries — limited source information (incl. 486 flagged unattested during the August 2026 description-research pass)
- All 3,668 creatures have source quality classifications.
Current source quality distribution (stories):
- Expert: 39 entries — reviewed by subject-matter specialists
- High: 55 entries — from high-quality references
- Medium: 351 entries — from intermediate references
- Good: 375 entries — from general references
- Fair: 795 entries — from general references or secondary sources
- Unknown: 570 entries — source pending classification
- All 2,185 stories have source quality classifications.
Versioning
The dataset is versioned with semantic versioning (MAJOR.MINOR.PATCH). Each version's date and change summary are recorded in the dataset manifest. The current version is displayed in entry footers. Major data updates trigger a new version release.
How to Contribute
If you have corrections, additions, or source materials to share:
- Email: mafolsson@tutamail.com
- Include the specific creature or story name, your source, and the correction
- For academic contributors: citations in MLA or Chicago style preferred
Limitations
- This is a living archive — completeness varies by region based on available source materials
- Some cultures are underrepresented due to limited published resources in English and Nordic languages
- The dataset is primarily text-based; illustrations are planned but not yet implemented
- Oral traditions are represented through published transcripts, which may not fully capture performance context
Last updated: 2026-08-03