Research queue

The lab's to-do list, in public

Project Aletheia is an independent lab that stress-tests contested science: take a claim, find the most boring explanation that could produce it, and run the test that decides between them. This page is everything the lab wants to test next — and, honestly, what each item still needs before it can run. Some need a dataset. Some need a better design. Some need a scientist with the right instrument. If you can supply the missing piece, any of these can move.

Experiments looking for a scientist

Fully designed experiments — question, protocol, budget, and the result that would kill them — waiting for a researcher with the right setup. The lab designs, you run, and the result publishes either way: a null here is a result, not a failure.

Package in preparationEXP-2026-001-RF-PULSE-AUDITIONPre-Foundry

Can simple radio-pulse patterns be heard and decoded?

Old military-era reports say people can hear certain radio-frequency pulses inside their head (a real, accepted effect called radio-frequency hearing) and one thin historical claim says simple pulse CODES could be interpreted. Nobody ever published the numbers. This experiment tests whether volunteers can decode simple pulse patterns above chance under modern safety limits.

Why it matters

It would replace a 60-year-old undocumented claim with the first real error-rate table for pulse-code radio hearing, and a clean null would finally bound what the historical record actually supports.

What it needs

  • RF exposure lab with safety (SAR) monitoring and human-subjects approval
  • pulse-modulated microwave source in the published RF-hearing range
  • sound-shielded test room
  • 10-20 adult volunteers

Budget

Gold

$100k+ standalone with dedicated exposure hardware and full board review

Shoestring

not runnable on a shoestring: RF human exposure requires a properly equipped lab

Standard

$15k-40k as a hosted study inside an existing RF/bioelectromagnetics lab (facility time, safety review, subject payments)

Time: 2-4 months including ethics review; sessions are under one hour per subject · Feasibility: HARD - needs a specialized RF exposure facility and full human-subjects review; realistic only as a collaboration with an existing bioelectromagnetics lab

Kill condition: Decoding accuracy indistinguishable from chance across the preregistered trial count at safe exposure levels

Claim ceiling: Bounded psychophysics only: whether simple pulse symbols are decodable above chance. Not speech, not communication devices, no clinical claims.

Source: Aletheia RF-ultrasound auditory investigation (2026-06); full protocol draft with safety preconditions exists and ships to the hosting lab

I could run this — talk to me
Seeking a scientistEXP-2026-002-VIBRATION-ONSETPre-Foundry

Do body-vibration episodes really come right before out-of-body experiences?

Thousands of people who report leaving their body describe a strong body-wide vibration just before it happens. All existing evidence is after-the-fact storytelling. This study has sleepers who often report such episodes log every vibration event prospectively (diary plus a simple wearable) so the timing claim gets tested forward instead of remembered backward.

Why it matters

The vibration-first sequence is one of the most consistent patterns in a century of these reports (Aletheia's own text analysis of 2,192 accounts found it splits cleanly across independent halves of the data). Nobody has ever tested it prospectively. Either outcome is a real result about how these experiences work.

What it needs

  • a sleep researcher or consciousness lab willing to host
  • 20-40 frequent experiencers (recruitable through existing experiencer communities)
  • consumer sleep wearables (accelerometer plus heart rate)
  • a preregistered diary protocol

Budget

Gold

$60k+ with in-lab polysomnography nights for a subset

Shoestring

$3k-6k fully remote: wearables shipped to participants, online diaries, subject payments

Standard

$10k-25k with a hosting lab, better wearables, and compliance monitoring

Time: 3-6 months of logging; setup one month; analysis is prewritten · Feasibility: MEDIUM - fully remote shoestring version is genuinely runnable; the hard part is disciplined recruitment, not equipment

Kill condition: Prospective logs show vibration episodes are NOT preferentially followed by the experiences within the preregistered window

Claim ceiling: Timing and sequence of self-reported episodes only. Says nothing about what the experiences ARE - only about whether the reported sequence survives prospective logging.

Source: Aletheia vibration-transition program (2026): corpus analyses, a full grant-grade study design, and the prewritten analysis plan exist

I could run this — talk to me
Validated package · resource blockedEXP-2026-003-LIGHT-TOUCH-HRVFoundry validated

Light Touch 30-minute HRV showdown

Does 30 minutes of Light Touch concha stimulation change paired log RMSSD more than a sensation-matched sham in healthy adults?

Why it matters

Give a qualified researcher a complete, pre-tested protocol for one narrow short-term HRV comparison.

What it needs

  • Light Touch active and sensation-matched sham configurations
  • Beat-level RR recording with Polar H10 or laboratory ECG
  • Recruit 48 healthy adults to retain 40 complete pairs

Budget

Gold

$16092.84 cash; 210 person-hours; 90 calendar days

Shoestring

$8773.96 cash; 110 person-hours; 90 calendar days

Standard

$10025.84 cash; 150 person-hours; 75 calendar days

Time: Standard tier: 150 person-hours across 75 calendar days. · Feasibility: Grade D (44.5/100) from the frozen cost, calendar, instrument, and recruitment formula.

Kill condition: Kill the positive-effect route if 40 valid pairs do not produce a positive mean with two-sided p below 0.05.

Claim ceiling: At most, this estimates one short-term paired RMSSD contrast in healthy adults under the exact tested dose and sham.

Source: /Users/bo/data/adp/investigations/experiment_foundry_plan_2026-08-24/known_answer_light_touch/FINDING_EXPANSION.json

I could run this — talk to me

The question queue

Every open question the lab has vetted, with its status. Each card expands to show the honest state of its test: what data exists, what is blocking it, the boring explanation that has to be ruled out first, and the most the test could ever prove even if everything goes right.

New here? What the labels mean
Ready now —
the data, the checks, and the kill rule all exist; this could run today.
Needs data —
the question is solid but a dataset is missing. If you have it, that is the whole blocker.
Needs design —
we do not yet have a test that could not fool itself.
Needs a collaborator —
requires an instrument, lab, or subject pool the lab does not have.
“The most this could ever prove” —
every test states its ceiling up front, so a narrow result can never quietly inflate into a big claim.
Needs designQ-COMMUNICATION-SPLIT

Can a computer separate reports with different kinds of communication when it is tested on new examples checked by people? The first pattern may only reflect how each source was written.

Precise form: What survives when the exploratory communication split pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-COMMUNICATION-SPLIT-V1

Needs designQ-DISCOURSE-QUALITY

Can a computer separate the quality and structure of different reports when it is tested on new examples checked by people? The first pattern may only reflect writing style or source.

Precise form: What survives when the exploratory discourse quality pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-DISCOURSE-QUALITY-V1

Needs designQ-DREAM-EXPERIENCE-EEG

Does the brain signal linked to reporting a dream reflect dreaming itself, or only lighter sleep and greater alertness? Separating these would show whether the signal is specific.

Precise form: Is the dream-report EEG correlate more than arousal or sleep depth?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Archive data and exploratory result exist
What's stopping us
analysis design and adequate raw channels
The boring explanation we must rule out first
stage, time, absolute power, and arousal
How we'd check it honestly
Stage/time-matched replication with absolute power and source localization
What would kill it
Kill consciousness-specific framing if arousal/depth controls absorb the effect.
The most this could ever prove
Archive-local physiological correlate only.

TEST-DREAM-EXPERIENCE-EEG-V1

Needs designQ-EM-MEASUREMENT-INVARIANCE

Do labels for unusual electrical or magnetic events mean the same thing across different report collections? If not, comparisons between collections may be artifacts of wording.

Precise form: What survives when the exploratory em measurement invariance pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-EM-MEASUREMENT-INVARIANCE-V1

Needs designQ-ENVIRONMENTAL-SUBSTRATE

Are unusual trickster-like reports more common in wet, cave, and forest settings after blind review and fair source controls? This tests whether the pattern is environmental or only a feature of the stories.

Precise form: Does the wet/cave/forest lower-trickster pattern survive blinded row adjudication and corpus controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Existing 120-row stratified packet
What's stopping us
coding rubric and second coder
The boring explanation we must rule out first
genre words, setting mentions, report length, and source channel
How we'd check it honestly
Two blinded human coders plus held-out adjudication agreement and within-corpus estimates
What would kill it
Kill if adjudicated positives collapse to lexical artifacts or within-corpus effects vanish.
The most this could ever prove
One packet/corpus ecology result; no environmental causation.

TEST-ENVIRONMENTAL-SUBSTRATE-V1

Needs designQ-LIGHT-SEPARATION

Can a computer separate reports about unusual lights from other reports when it is tested on new examples checked by people? The first pattern may only reflect source vocabulary.

Precise form: What survives when the exploratory light separation pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-LIGHT-SEPARATION-V1

Needs designQ-NDE-DREAM-MEANING

Do near-death reports and dream reports differ in meaning and communication after writing style, report length, and repeated authors are controlled? This tests whether the contrast belongs to the experiences or only to the archives.

Precise form: Do NDE and DreamBank meaning/communication contrasts survive source, length, and ontology controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
NDE and DreamBank text exist
What's stopping us
validated refined ontology
The boring explanation we must rule out first
source style, length, repeated people, and ontology leakage
How we'd check it honestly
Blinded codebook labels with held-out source-family tests
What would kill it
Kill if the contrast vanishes under source/length controls.
The most this could ever prove
Two-archive semantic contrast only.

TEST-NDE-DREAM-MEANING-V1

Needs designQ-ONSET-GRAMMAR

Do key events appear in a repeatable order across different experience reports, and is that order different from fiction? This tests whether the sequence is real or only a keyword and storytelling effect.

Precise form: Do validated hallmark events have reproducible onset order across experiential corpora and fiction controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
OBERF/NDERF and fiction text exist
What's stopping us
context-aware extractor and gold labels
The boring explanation we must rule out first
broad keyword position masquerading as event order
How we'd check it honestly
Blinded event-identity precision gate before any order statistic
What would kill it
Kill if event precision fails or order matches fiction/source artifacts.
The most this could ever prove
Validated corpus order only; no universal experiential sequence.

TEST-ONSET-GRAMMAR-V1

Needs designQ-ONTOLOGY-SAFETY

Can broad labels for electrical, magnetic, and time-related experiences be split into clearer meanings without losing useful distinctions? Better labels would reduce false patterns caused by words with several meanings.

Precise form: Can broad EM/time labels be replaced without losing valid distinctions?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Old labels and source-grounded rows exist
What's stopping us
gold annotations
The boring explanation we must rule out first
polysemy and source-specific word meaning
How we'd check it honestly
Manual adjudication benchmark comparing broad and split ontologies
What would kill it
Reject any replacement that fails held-out meaning discrimination.
The most this could ever prove
Ontology-performance result only.

TEST-ONTOLOGY-SAFETY-V1

Needs designQ-SENSORY-REDUCTION

Does the same pattern appear across reduced-sensory experiments, hypnosis, dreams, and sleep paralysis after the collections are fairly matched? This would show whether it is specific to sensory reduction or common to many report types.

Precise form: Does a sensory-reduction pattern survive matched Ganzfeld, hypnosis, dream, and sleep-paralysis controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Several local corpora; protocol coverage uneven
What's stopping us
shared coding instrument
The boring explanation we must rule out first
report purpose, length, and vocabulary leakage
How we'd check it honestly
Matched human-validated labels with corpus-held-out contrasts
What would kill it
Kill if controls reproduce the same pattern after matching.
The most this could ever prove
Corpus separation only; no sensory-gating mechanism.

TEST-SENSORY-REDUCTION-V1

Needs designQ-THREAT-PRESENCE-INVARIANCE

Do labels for threat and a sensed presence mean and perform the same way across different report collections? If not, comparisons may reflect genre and ordinary fear words.

Precise form: Are threat/presence labels measurement-invariant across source families?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Candidate corpora and proposed benchmark exist
What's stopping us
gold set and label rubric
The boring explanation we must rule out first
ordinary fear, entity nouns, genre conventions
How we'd check it honestly
Blinded multi-source gold set with per-source precision/recall floors
What would kill it
Kill invariance if accuracy or meaning shifts materially by source.
The most this could ever prove
Measurement result only; no entity interpretation.

TEST-THREAT-PRESENCE-INVARIANCE-V1

Needs designQ-TIME-MEASUREMENT-INVARIANCE

Do labels for unusual experiences of time mean the same thing across different report collections? If they change with source or wording, any cross-source pattern is unreliable.

Precise form: What survives when the exploratory time measurement invariance pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-TIME-MEASUREMENT-INVARIANCE-V1

Needs designQ-UAP-CRYPTID-ECOLOGY

Does the apparent link between reports of unexplained objects and unusual creatures survive new examples checked by people and fair source controls? Shared vocabulary may create the pattern.

Precise form: What survives when the exploratory uap cryptid ecology pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-UAP-CRYPTID-ECOLOGY-V1

Needs designQ-UAP-EVIDENCE-CHANNEL

Can reports be separated by the kind of evidence they contain when tested on new examples checked by people? The first split may only reflect how each source asks people to report.

Precise form: What survives when the exploratory uap evidence channel pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-UAP-EVIDENCE-CHANNEL-V1

Needs designQ-UAP-REPORTING-ECOLOGY

Do different channels for reporting unexplained objects produce repeatable differences after source and writing style are controlled? This would show whether the pattern belongs to the reporting systems rather than the events.

Precise form: What survives when the exploratory uap reporting ecology pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-UAP-REPORTING-ECOLOGY-V1

Needs designQ-OBERF-TELLING-ORDER

Do people tell the main events in near-death reports in the same order in a second, independently collected archive? A match would support a reporting-order pattern, not the order in which the events were experienced.

Precise form: Does the coarse telling-order gradient replicate in the independently collected NDERF corpus?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Local NDERF exports (nderf_experiences.csv, nderf_reports.csv)
What's stopping us
second precision rater
The boring explanation we must rule out first
narrative prompts and questionnaire structure shaping telling order
How we'd check it honestly
Frozen extractor v2 + precision gate with a second rater + event-set-preserving permutation null
What would kill it
Kill if either endpoint fails p<=0.005 in both NDERF splits with gate-passing markers.
The most this could ever prove
Telling order in one additional corpus; still not experienced order.

TEST-OBERF-TELLING-ORDER-V1

Needs designQ-GANZFELD-DB-STRUCTURE

Does a second collection of telepathy experiments show the same overall excess of correct choices, no decline over time, and only a weak small-study pattern after duplicate studies are removed? This tests the compiled tables, not whether telepathy is real.

Precise form: Does the database-level structure (robust pooled excess, no decline, weak small-study signal) hold in the independent Tressoldi compilation?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Local MA_GanzfeldESP.xlsx (Tressoldi set)
What's stopping us
none
The boring explanation we must rule out first
overlapping study membership between compilations
How we'd check it honestly
Same frozen audit statistics plus a study-overlap dedup join between the two files
What would kill it
Kill consistency if conclusions flip after overlap dedup at frozen thresholds.
The most this could ever prove
Structure of compiled databases; upstream selection untested.

TEST-GANZFELD-DB-STRUCTURE-V1

Looking for what already ran? Results live on Findings. Investigations that ended without an answer — and what would revive them — live on Cold Cases.