Every open question the lab has vetted, with its status. Each card expands to show the honest state of its test: what data exists, what is blocking it, the boring explanation that has to be ruled out first, and the most the test could ever prove even if everything goes right.
New here? What the labels mean
- Ready now —
- the data, the checks, and the kill rule all exist; this could run today.
- Needs data —
- the question is solid but a dataset is missing. If you have it, that is the whole blocker.
- Needs design —
- we do not yet have a test that could not fool itself.
- Needs a collaborator —
- requires an instrument, lab, or subject pool the lab does not have.
- “The most this could ever prove” —
- every test states its ceiling up front, so a narrow result can never quietly inflate into a big claim.
Needs designQ-COMMUNICATION-SPLIT
Can a computer separate reports with different kinds of communication when it is tested on new examples checked by people? The first pattern may only reflect how each source was written.
Precise form: What survives when the exploratory communication split pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-COMMUNICATION-SPLIT-V1
Needs designQ-DISCOURSE-QUALITY
Can a computer separate the quality and structure of different reports when it is tested on new examples checked by people? The first pattern may only reflect writing style or source.
Precise form: What survives when the exploratory discourse quality pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-DISCOURSE-QUALITY-V1
Needs designQ-DREAM-EXPERIENCE-EEG
Does the brain signal linked to reporting a dream reflect dreaming itself, or only lighter sleep and greater alertness? Separating these would show whether the signal is specific.
Precise form: Is the dream-report EEG correlate more than arousal or sleep depth?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Archive data and exploratory result exist
- What's stopping us
- analysis design and adequate raw channels
- The boring explanation we must rule out first
- stage, time, absolute power, and arousal
- How we'd check it honestly
- Stage/time-matched replication with absolute power and source localization
- What would kill it
- Kill consciousness-specific framing if arousal/depth controls absorb the effect.
- The most this could ever prove
- Archive-local physiological correlate only.
TEST-DREAM-EXPERIENCE-EEG-V1
Needs designQ-EM-MEASUREMENT-INVARIANCE
Do labels for unusual electrical or magnetic events mean the same thing across different report collections? If not, comparisons between collections may be artifacts of wording.
Precise form: What survives when the exploratory em measurement invariance pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-EM-MEASUREMENT-INVARIANCE-V1
Needs designQ-ENVIRONMENTAL-SUBSTRATE
Are unusual trickster-like reports more common in wet, cave, and forest settings after blind review and fair source controls? This tests whether the pattern is environmental or only a feature of the stories.
Precise form: Does the wet/cave/forest lower-trickster pattern survive blinded row adjudication and corpus controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Existing 120-row stratified packet
- What's stopping us
- coding rubric and second coder
- The boring explanation we must rule out first
- genre words, setting mentions, report length, and source channel
- How we'd check it honestly
- Two blinded human coders plus held-out adjudication agreement and within-corpus estimates
- What would kill it
- Kill if adjudicated positives collapse to lexical artifacts or within-corpus effects vanish.
- The most this could ever prove
- One packet/corpus ecology result; no environmental causation.
TEST-ENVIRONMENTAL-SUBSTRATE-V1
Needs designQ-LIGHT-SEPARATION
Can a computer separate reports about unusual lights from other reports when it is tested on new examples checked by people? The first pattern may only reflect source vocabulary.
Precise form: What survives when the exploratory light separation pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-LIGHT-SEPARATION-V1
Needs designQ-NDE-DREAM-MEANING
Do near-death reports and dream reports differ in meaning and communication after writing style, report length, and repeated authors are controlled? This tests whether the contrast belongs to the experiences or only to the archives.
Precise form: Do NDE and DreamBank meaning/communication contrasts survive source, length, and ontology controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- NDE and DreamBank text exist
- What's stopping us
- validated refined ontology
- The boring explanation we must rule out first
- source style, length, repeated people, and ontology leakage
- How we'd check it honestly
- Blinded codebook labels with held-out source-family tests
- What would kill it
- Kill if the contrast vanishes under source/length controls.
- The most this could ever prove
- Two-archive semantic contrast only.
TEST-NDE-DREAM-MEANING-V1
Needs designQ-ONSET-GRAMMAR
Do key events appear in a repeatable order across different experience reports, and is that order different from fiction? This tests whether the sequence is real or only a keyword and storytelling effect.
Precise form: Do validated hallmark events have reproducible onset order across experiential corpora and fiction controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- OBERF/NDERF and fiction text exist
- What's stopping us
- context-aware extractor and gold labels
- The boring explanation we must rule out first
- broad keyword position masquerading as event order
- How we'd check it honestly
- Blinded event-identity precision gate before any order statistic
- What would kill it
- Kill if event precision fails or order matches fiction/source artifacts.
- The most this could ever prove
- Validated corpus order only; no universal experiential sequence.
TEST-ONSET-GRAMMAR-V1
Needs designQ-ONTOLOGY-SAFETY
Can broad labels for electrical, magnetic, and time-related experiences be split into clearer meanings without losing useful distinctions? Better labels would reduce false patterns caused by words with several meanings.
Precise form: Can broad EM/time labels be replaced without losing valid distinctions?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Old labels and source-grounded rows exist
- What's stopping us
- gold annotations
- The boring explanation we must rule out first
- polysemy and source-specific word meaning
- How we'd check it honestly
- Manual adjudication benchmark comparing broad and split ontologies
- What would kill it
- Reject any replacement that fails held-out meaning discrimination.
- The most this could ever prove
- Ontology-performance result only.
TEST-ONTOLOGY-SAFETY-V1
Needs designQ-SENSORY-REDUCTION
Does the same pattern appear across reduced-sensory experiments, hypnosis, dreams, and sleep paralysis after the collections are fairly matched? This would show whether it is specific to sensory reduction or common to many report types.
Precise form: Does a sensory-reduction pattern survive matched Ganzfeld, hypnosis, dream, and sleep-paralysis controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Several local corpora; protocol coverage uneven
- What's stopping us
- shared coding instrument
- The boring explanation we must rule out first
- report purpose, length, and vocabulary leakage
- How we'd check it honestly
- Matched human-validated labels with corpus-held-out contrasts
- What would kill it
- Kill if controls reproduce the same pattern after matching.
- The most this could ever prove
- Corpus separation only; no sensory-gating mechanism.
TEST-SENSORY-REDUCTION-V1
Needs designQ-THREAT-PRESENCE-INVARIANCE
Do labels for threat and a sensed presence mean and perform the same way across different report collections? If not, comparisons may reflect genre and ordinary fear words.
Precise form: Are threat/presence labels measurement-invariant across source families?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Candidate corpora and proposed benchmark exist
- What's stopping us
- gold set and label rubric
- The boring explanation we must rule out first
- ordinary fear, entity nouns, genre conventions
- How we'd check it honestly
- Blinded multi-source gold set with per-source precision/recall floors
- What would kill it
- Kill invariance if accuracy or meaning shifts materially by source.
- The most this could ever prove
- Measurement result only; no entity interpretation.
TEST-THREAT-PRESENCE-INVARIANCE-V1
Needs designQ-TIME-MEASUREMENT-INVARIANCE
Do labels for unusual experiences of time mean the same thing across different report collections? If they change with source or wording, any cross-source pattern is unreliable.
Precise form: What survives when the exploratory time measurement invariance pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-TIME-MEASUREMENT-INVARIANCE-V1
Needs designQ-UAP-CRYPTID-ECOLOGY
Does the apparent link between reports of unexplained objects and unusual creatures survive new examples checked by people and fair source controls? Shared vocabulary may create the pattern.
Precise form: What survives when the exploratory uap cryptid ecology pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-UAP-CRYPTID-ECOLOGY-V1
Needs designQ-UAP-EVIDENCE-CHANNEL
Can reports be separated by the kind of evidence they contain when tested on new examples checked by people? The first split may only reflect how each source asks people to report.
Precise form: What survives when the exploratory uap evidence channel pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-UAP-EVIDENCE-CHANNEL-V1
Needs designQ-UAP-REPORTING-ECOLOGY
Do different channels for reporting unexplained objects produce repeatable differences after source and writing style are controlled? This would show whether the pattern belongs to the reporting systems rather than the events.
Precise form: What survives when the exploratory uap reporting ecology pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-UAP-REPORTING-ECOLOGY-V1
Needs designQ-OBERF-TELLING-ORDER
Do people tell the main events in near-death reports in the same order in a second, independently collected archive? A match would support a reporting-order pattern, not the order in which the events were experienced.
Precise form: Does the coarse telling-order gradient replicate in the independently collected NDERF corpus?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Local NDERF exports (nderf_experiences.csv, nderf_reports.csv)
- What's stopping us
- second precision rater
- The boring explanation we must rule out first
- narrative prompts and questionnaire structure shaping telling order
- How we'd check it honestly
- Frozen extractor v2 + precision gate with a second rater + event-set-preserving permutation null
- What would kill it
- Kill if either endpoint fails p<=0.005 in both NDERF splits with gate-passing markers.
- The most this could ever prove
- Telling order in one additional corpus; still not experienced order.
TEST-OBERF-TELLING-ORDER-V1
Needs designQ-GANZFELD-DB-STRUCTURE
Does a second collection of telepathy experiments show the same overall excess of correct choices, no decline over time, and only a weak small-study pattern after duplicate studies are removed? This tests the compiled tables, not whether telepathy is real.
Precise form: Does the database-level structure (robust pooled excess, no decline, weak small-study signal) hold in the independent Tressoldi compilation?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Local MA_GanzfeldESP.xlsx (Tressoldi set)
- What's stopping us
- none
- The boring explanation we must rule out first
- overlapping study membership between compilations
- How we'd check it honestly
- Same frozen audit statistics plus a study-overlap dedup join between the two files
- What would kill it
- Kill consistency if conclusions flip after overlap dedup at frozen thresholds.
- The most this could ever prove
- Structure of compiled databases; upstream selection untested.
TEST-GANZFELD-DB-STRUCTURE-V1