Supported findingVerified within one archive · 2026-07-23

De-identified dream reports can be linked across time

A computer often linked a person’s later DreamBank reports to their earlier reports using only a fixed list of common grammar and discourse words. Names, places, occupations, topic terms, and report length were removed as clues.

32 of 100
Correct first match
5 of 100 by chance
51 of 100
Correct within three matches
15 of 100 by chance
12 of 20
Contributors linked above chance
8 did not clear chance

What the investigation tested

Can grammar reconnect a record to its contributor?

We used 20 individual DreamBank contributors. The model learned from the first 80% of each person’s ordered reports, then tried to link the final 20% back to the correct earlier contributor record.

The privacy mask retained only 161 common grammar and discourse words, such as and, because, with, and although. Names, nouns, places, occupations, and topic terms were excluded. Each report was also normalized so simple verbosity could not drive the result.

Balanced accuracy was 31.9%, compared with 5% chance. The correct contributor appeared among the first three matches 51.0% of the time, compared with 15% chance. Both permutation tests gave p=0.000999.

The result was not universal

Twelve of 20 contributors linked above top-match chance. Eight did not, and six had zero correct first matches. The signal is real at the group level but highly uneven across people.

The stronger idea failed

Waking prose did not reliably identify dream authors

A second frozen test trained on ordinary waking prose from five contributors and tried to identify their dream reports. Grammar reached 25.8% balanced accuracy against 20% chance; scrubbed vocabulary reached 25.5%. Neither survived the correction and effect-size gates.

That null changes the interpretation. The current result looks more like a stable dream-reporting or collection signature than a universal personal fingerprint that crosses from waking thought into dreams.

Scientific value

Why this matters

A privacy warning

Removing names and obvious personal details may not prevent longitudinal narrative records from being linked. Grammar and reporting habits can remain identifying.

A machine-learning warning

Randomly splitting reports by row can let a model learn the contributor instead of the phenomenon. Longitudinal text should be split by person or collection unit.

A boundary on the science

A separate waking-writing test failed. The evidence supports archive linkage, not a universal language fingerprint that passes from waking thought into dreams.

What it does not show

  • It does not recover anyone’s legal identity.
  • It does not show that dreams reveal biological identity.
  • It does not generalize beyond DreamBank yet.
  • It does not separate the person from recording or transcription method.

The decisive next test

Use a new cohort with independently collected waking and dream writing, separate recorder and transcriber identities, and a true open-set test that can answer “unknown person.” That would show whether the signal belongs to a person, a recording method, or this archive.

How we checked it

  • Both analyses were preregistered and SHA-256 frozen before results.
  • Known-answer positive and shuffled-label controls passed.
  • Privacy endpoints used 1,000 frozen label permutations.
  • A separate sklearn implementation reproduced every headline number to 12 decimal places.

Source: DreamBank, Adam Schneider and G. William Domhoff, UC Santa Cruz. DreamBank data are licensed CC BY-NC-SA 4.0.