Research queue

The lab's to-do list, in public

Project Aletheia is an independent lab that stress-tests contested science: take a claim, find the most boring explanation that could produce it, and run the test that decides between them. This page is everything the lab wants to test next — and, honestly, what each item still needs before it can run. Some need a dataset. Some need a better design. Some need a scientist with the right instrument. If you can supply the missing piece, any of these can move.

Experiments looking for a scientist

Fully designed experiments — question, protocol, budget, and the result that would kill them — waiting for a researcher with the right setup. The lab designs, you run, and the result publishes either way: a null here is a result, not a failure.

Package in preparationEXP-2026-001-RF-PULSE-AUDITIONPre-Foundry

Can simple radio-pulse patterns be heard and decoded?

Old military-era reports say people can hear certain radio-frequency pulses inside their head (a real, accepted effect called radio-frequency hearing) and one thin historical claim says simple pulse CODES could be interpreted. Nobody ever published the numbers. This experiment tests whether volunteers can decode simple pulse patterns above chance under modern safety limits.

Why it matters

It would replace a 60-year-old undocumented claim with the first real error-rate table for pulse-code radio hearing, and a clean null would finally bound what the historical record actually supports.

What it needs

  • RF exposure lab with safety (SAR) monitoring and human-subjects approval
  • pulse-modulated microwave source in the published RF-hearing range
  • sound-shielded test room
  • 10-20 adult volunteers

Budget

Gold

$100k+ standalone with dedicated exposure hardware and full board review

Shoestring

not runnable on a shoestring: RF human exposure requires a properly equipped lab

Standard

$15k-40k as a hosted study inside an existing RF/bioelectromagnetics lab (facility time, safety review, subject payments)

Time: 2-4 months including ethics review; sessions are under one hour per subject · Feasibility: HARD - needs a specialized RF exposure facility and full human-subjects review; realistic only as a collaboration with an existing bioelectromagnetics lab

Kill condition: Decoding accuracy indistinguishable from chance across the preregistered trial count at safe exposure levels

Claim ceiling: Bounded psychophysics only: whether simple pulse symbols are decodable above chance. Not speech, not communication devices, no clinical claims.

Source: Aletheia RF-ultrasound auditory investigation (2026-06); full protocol draft with safety preconditions exists and ships to the hosting lab

I could run this — talk to me
Seeking a scientistEXP-2026-002-VIBRATION-ONSETPre-Foundry

Do body-vibration episodes really come right before out-of-body experiences?

Thousands of people who report leaving their body describe a strong body-wide vibration just before it happens. All existing evidence is after-the-fact storytelling. This study has sleepers who often report such episodes log every vibration event prospectively (diary plus a simple wearable) so the timing claim gets tested forward instead of remembered backward.

Why it matters

The vibration-first sequence is one of the most consistent patterns in a century of these reports (Aletheia's own text analysis of 2,192 accounts found it splits cleanly across independent halves of the data). Nobody has ever tested it prospectively. Either outcome is a real result about how these experiences work.

What it needs

  • a sleep researcher or consciousness lab willing to host
  • 20-40 frequent experiencers (recruitable through existing experiencer communities)
  • consumer sleep wearables (accelerometer plus heart rate)
  • a preregistered diary protocol

Budget

Gold

$60k+ with in-lab polysomnography nights for a subset

Shoestring

$3k-6k fully remote: wearables shipped to participants, online diaries, subject payments

Standard

$10k-25k with a hosting lab, better wearables, and compliance monitoring

Time: 3-6 months of logging; setup one month; analysis is prewritten · Feasibility: MEDIUM - fully remote shoestring version is genuinely runnable; the hard part is disciplined recruitment, not equipment

Kill condition: Prospective logs show vibration episodes are NOT preferentially followed by the experiences within the preregistered window

Claim ceiling: Timing and sequence of self-reported episodes only. Says nothing about what the experiences ARE - only about whether the reported sequence survives prospective logging.

Source: Aletheia vibration-transition program (2026): corpus analyses, a full grant-grade study design, and the prewritten analysis plan exist

I could run this — talk to me
Validated package · resource blockedEXP-2026-003-LIGHT-TOUCH-HRVFoundry validated

Light Touch 30-minute HRV showdown

Does 30 minutes of Light Touch concha stimulation change paired log RMSSD more than a sensation-matched sham in healthy adults?

Why it matters

Give a qualified researcher a complete, pre-tested protocol for one narrow short-term HRV comparison.

What it needs

  • Light Touch active and sensation-matched sham configurations
  • Beat-level RR recording with Polar H10 or laboratory ECG
  • Recruit 48 healthy adults to retain 40 complete pairs

Budget

Gold

$16092.84 cash; 210 person-hours; 90 calendar days

Shoestring

$8773.96 cash; 110 person-hours; 90 calendar days

Standard

$10025.84 cash; 150 person-hours; 75 calendar days

Time: Standard tier: 150 person-hours across 75 calendar days. · Feasibility: Grade D (44.5/100) from the frozen cost, calendar, instrument, and recruitment formula.

Kill condition: Kill the positive-effect route if 40 valid pairs do not produce a positive mean with two-sided p below 0.05.

Claim ceiling: At most, this estimates one short-term paired RMSSD contrast in healthy adults under the exact tested dose and sham.

Source: /Users/bo/data/adp/investigations/experiment_foundry_plan_2026-08-24/known_answer_light_touch/FINDING_EXPANSION.json

I could run this — talk to me

The question queue

Every open question the lab has vetted, with its status. Each card expands to show the honest state of its test: what data exists, what is blocking it, the boring explanation that has to be ruled out first, and the most the test could ever prove even if everything goes right.

New here? What the labels mean
Ready now —
the data, the checks, and the kill rule all exist; this could run today.
Needs data —
the question is solid but a dataset is missing. If you have it, that is the whole blocker.
Needs design —
we do not yet have a test that could not fool itself.
Needs a collaborator —
requires an instrument, lab, or subject pool the lab does not have.
“The most this could ever prove” —
every test states its ceiling up front, so a narrow result can never quietly inflate into a big claim.
Needs dataQ-BIOLOGY-RECEPTOR

Do the 27 genes really form a biological link, or did name changes and mixed gene families create it? A false link could send later research in the wrong direction.

Precise form: Do receptor-family bridges survive exact 27-gene identity and provenance checks?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Existing 27-gene ledgers; expression and true receptor-family universes are absent
What's stopping us
expression matrix and receptor-family annotations
The boring explanation we must rule out first
gene-symbol aliases and mixed receptor families
How we'd check it honestly
Deterministic provenance join plus expression/family-matched nulls with planted controls
What would kill it
Kill any bridge whose members lack exact source-backed receptor-family identity.
The most this could ever prove
Dataset-specific provenance result; no biological mechanism.

TEST-BIOLOGY-RECEPTOR-V1

Needs dataQ-BODY-BOUNDARY

Do fainting, coma, anesthesia, sleep paralysis, near-death experiences, and extreme acceleration produce different patterns of feeling outside the body? Fair samples could show whether experiences that sound alike have different triggers.

Precise form: Do G-LOC, syncope, coma, anesthesia, sleep-paralysis, and NDE data preserve trigger-specific body-boundary patterns?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Some coded narratives; raw denominators incomplete
What's stopping us
raw denominators and independent labels
The boring explanation we must rule out first
keyword extraction, selected case series, unequal denominators
How we'd check it honestly
Blinded precision audit followed by diagnosis-matched proportion contrasts
What would kill it
Kill if validated labels or non-selected denominators remove the trigger contrast.
The most this could ever prove
Trigger/corpus contrast only; no shared mechanism.

TEST-BODY-BOUNDARY-V1

Needs designQ-COMMUNICATION-SPLIT

Can a computer separate reports with different kinds of communication when it is tested on new examples checked by people? The first pattern may only reflect how each source was written.

Precise form: What survives when the exploratory communication split pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-COMMUNICATION-SPLIT-V1

Needs dataQ-CORT-INTERMISSION

Do reports of an interval between lives form a repeatable pattern when every detail is traced to the original case record? This would show whether the proposed subtype is more than a set of memorable examples.

Precise form: Do intermission features survive a source-traceable case-level matrix?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Public CORT sources exist; case matrix incomplete
What's stopping us
case-level extraction
The boring explanation we must rule out first
secondary summaries and selected memorable cases
How we'd check it honestly
Double-entry source verification with blind feature coding
What would kill it
Kill if source-verifiable case rates do not reproduce the proposed subtype.
The most this could ever prove
Documented-case pattern only.

TEST-CORT-INTERMISSION-V1

Needs a collaboratorQ-CORT-PROSPECTIVE

Can claims about a previous life be recorded before anyone searches for a matching family, then match better than decoys in a blind test? This would test whether information leaks explain the apparent matches.

Precise form: Can prospective reincarnation-case documentation beat family and investigator information leakage?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Historical cases only
What's stopping us
participants and field collaborators
The boring explanation we must rule out first
retrospective memory, family cueing, investigator leakage
How we'd check it honestly
Prospective sealed statements, independent matching, and blinded scoring
What would kill it
Kill if blinded matches do not exceed preregistered decoys.
The most this could ever prove
Prospective case documentation only.

EXP-02

Needs a collaboratorQ-CRISIS-APPARITION-PROSPECTIVE

If people record unusual appearances when they happen, do those reports coincide with a death more often than chance and reporting opportunity allow? A modern test could avoid the memory and selection problems in old stories.

Precise form: Does a modern contemporaneous census show death-time coincidence beyond reporting opportunity?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Historical literature; no modern census
What's stopping us
participants, ethics, outcome linkage
The boring explanation we must rule out first
retrospective dating, denominator and reporting bias
How we'd check it honestly
Prospective timestamped reports linked blindly to outcome records
What would kill it
Kill if coincidence does not exceed the frozen opportunity-adjusted null.
The most this could ever prove
Prospective cohort association only.

EXP-03

Needs designQ-DISCOURSE-QUALITY

Can a computer separate the quality and structure of different reports when it is tested on new examples checked by people? The first pattern may only reflect writing style or source.

Precise form: What survives when the exploratory discourse quality pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-DISCOURSE-QUALITY-V1

Needs a collaboratorQ-DMT-NDE-BRIDGE

Do reports from the drug DMT and near-death experiences share anything specific after common words and general altered-state features are removed? This would test broad claims that the two experiences share one pattern.

Precise form: Do DMT and NDE reports share a specific signal beyond vocabulary and generic altered-state content?

4 linked tests · last triaged 2026-07-23

The honest state of this test
What data we already have
Low-density/archive data tested; decisive paired data absent
What's stopping us
new high-density paired datasets
The boring explanation we must rule out first
item wording, vocabulary leakage, resolution mismatch
How we'd check it honestly
Held-out validated coding; physiology requires preregistered paired high-density data
What would kill it
Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
The most this could ever prove
Dataset-specific bridge or bounded null; no common mechanism.

EXP-BRIDGE-25-01

The honest state of this test
What data we already have
Low-density/archive data tested; decisive paired data absent
What's stopping us
new high-density paired datasets
The boring explanation we must rule out first
item wording, vocabulary leakage, resolution mismatch
How we'd check it honestly
Held-out validated coding; physiology requires preregistered paired high-density data
What would kill it
Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
The most this could ever prove
Dataset-specific bridge or bounded null; no common mechanism.

EXP-BRIDGE-25-02

The honest state of this test
What data we already have
Low-density/archive data tested; decisive paired data absent
What's stopping us
new high-density paired datasets
The boring explanation we must rule out first
item wording, vocabulary leakage, resolution mismatch
How we'd check it honestly
Held-out validated coding; physiology requires preregistered paired high-density data
What would kill it
Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
The most this could ever prove
Dataset-specific bridge or bounded null; no common mechanism.

EXP-BRIDGE-25-03

The honest state of this test
What data we already have
Low-density/archive data tested; decisive paired data absent
What's stopping us
new high-density paired datasets
The boring explanation we must rule out first
item wording, vocabulary leakage, resolution mismatch
How we'd check it honestly
Held-out validated coding; physiology requires preregistered paired high-density data
What would kill it
Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
The most this could ever prove
Dataset-specific bridge or bounded null; no common mechanism.

EXP-BRIDGE-25-04

Needs designQ-DREAM-EXPERIENCE-EEG

Does the brain signal linked to reporting a dream reflect dreaming itself, or only lighter sleep and greater alertness? Separating these would show whether the signal is specific.

Precise form: Is the dream-report EEG correlate more than arousal or sleep depth?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Archive data and exploratory result exist
What's stopping us
analysis design and adequate raw channels
The boring explanation we must rule out first
stage, time, absolute power, and arousal
How we'd check it honestly
Stage/time-matched replication with absolute power and source localization
What would kill it
Kill consciousness-specific framing if arousal/depth controls absorb the effect.
The most this could ever prove
Archive-local physiological correlate only.

TEST-DREAM-EXPERIENCE-EEG-V1

Needs dataQ-DREAM-PRIVACY-LINKAGE

Can dream reports be linked to the same writer in a new group when the possible author is not limited to a small known set? This matters because reliable linkage could create a privacy risk.

Precise form: Does dream-report linkage generalize to independent open-set contributors?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
DreamBank result only; external cohort absent
What's stopping us
independent consented cohort
The boring explanation we must rule out first
recurring content, recorder/transcriber, closed-set advantage
How we'd check it honestly
Open-set external cohort with contributor and recorder separation
What would kill it
Kill generalization if open-set linkage falls to the frozen null.
The most this could ever prove
Current finding remains one-archive closed-set linkage.

TEST-DREAM-PRIVACY-LINKAGE-V1

Needs designQ-EM-MEASUREMENT-INVARIANCE

Do labels for unusual electrical or magnetic events mean the same thing across different report collections? If not, comparisons between collections may be artifacts of wording.

Precise form: What survives when the exploratory em measurement invariance pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-EM-MEASUREMENT-INVARIANCE-V1

Needs designQ-ENVIRONMENTAL-SUBSTRATE

Are unusual trickster-like reports more common in wet, cave, and forest settings after blind review and fair source controls? This tests whether the pattern is environmental or only a feature of the stories.

Precise form: Does the wet/cave/forest lower-trickster pattern survive blinded row adjudication and corpus controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Existing 120-row stratified packet
What's stopping us
coding rubric and second coder
The boring explanation we must rule out first
genre words, setting mentions, report length, and source channel
How we'd check it honestly
Two blinded human coders plus held-out adjudication agreement and within-corpus estimates
What would kill it
Kill if adjudicated positives collapse to lexical artifacts or within-corpus effects vanish.
The most this could ever prove
One packet/corpus ecology result; no environmental causation.

TEST-ENVIRONMENTAL-SUBSTRATE-V1

Needs a collaboratorQ-GCP-PROSPECTIVE

During a future major event, does a worldwide network of random-number devices behave unusually after hardware faults and event selection are fixed in advance? A prospective test avoids choosing the event or analysis after seeing the data.

Precise form: Does a preregistered future event produce network behavior beyond hardware and selection controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Retrospective archive only
What's stopping us
future events and live network
The boring explanation we must rule out first
event selection, device drift, archive sensitivity choices
How we'd check it honestly
Prospective preregistration with hardware diagnostics and untouched endpoint
What would kill it
Kill if prospective endpoint stays inside its frozen null.
The most this could ever prove
Prospective network statistic only.

EXP-01

Needs dataQ-HESSDALEN-ECOLOGY

Do reported lights in Hessdalen still show the same links to weather and solar activity when direct station records and new dates are used? This would test the first result with better local data.

Precise form: Does the matched weather/Kp result replicate on direct station data and held-out dates?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Current dated corpus; direct station coverage incomplete
What's stopping us
station logs and independent dates
The boring explanation we must rule out first
reanalysis-grid error and observer opportunity
How we'd check it honestly
Held-out case-crossover using station logs and observer-density controls
What would kill it
Kill extension if Kp elevation appears or weather differences vanish.
The most this could ever prove
Current finding stays reanalysis-based observation ecology.

TEST-HESSDALEN-ECOLOGY-V1

Needs designQ-LIGHT-SEPARATION

Can a computer separate reports about unusual lights from other reports when it is tested on new examples checked by people? The first pattern may only reflect source vocabulary.

Precise form: What survives when the exploratory light separation pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-LIGHT-SEPARATION-V1

Needs dataQ-LZC-APERIODIC

After ordinary changes in the brain-wave spectrum are removed, does the signal-complexity measure still show no drug effect in new recordings? This would test whether the earlier null survives a cleaner measurement.

Precise form: Does the LZc effect remain absent after preregistered aperiodic control in new raw EEG?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Two retrospective datasets; prospective replication absent
What's stopping us
independent raw EEG
The boring explanation we must rule out first
spectral slope, power, state mixing, and small subjects
How we'd check it honestly
Person-level preregistered mixed model with spectral-surrogate comparison
What would kill it
Falsify the current claim if a powered drug effect survives slope control.
The most this could ever prove
Current claim remains two-dataset measurement result.

TEST-LZC-APERIODIC-V1

Needs designQ-NDE-DREAM-MEANING

Do near-death reports and dream reports differ in meaning and communication after writing style, report length, and repeated authors are controlled? This tests whether the contrast belongs to the experiences or only to the archives.

Precise form: Do NDE and DreamBank meaning/communication contrasts survive source, length, and ontology controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
NDE and DreamBank text exist
What's stopping us
validated refined ontology
The boring explanation we must rule out first
source style, length, repeated people, and ontology leakage
How we'd check it honestly
Blinded codebook labels with held-out source-family tests
What would kill it
Kill if the contrast vanishes under source/length controls.
The most this could ever prove
Two-archive semantic contrast only.

TEST-NDE-DREAM-MEANING-V1

Needs dataQ-NDE-LIFE-REVIEW

Can the reported link between life reviews and encounters with deceased relatives be reproduced from checked, row-level near-death reports? The original percentages cannot yet be traced to a validated table.

Precise form: Is the reported life-review/deceased-relative subtype reproducible from validated NDE rows?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Claim summary exists; validated row packet not found
What's stopping us
recover coded rows
The boring explanation we must rule out first
generated labels and denominator loss
How we'd check it honestly
Blind label validation plus exact denominator recomputation
What would kill it
Kill if 58.14% and 16.279% cannot be reproduced on validated labels.
The most this could ever prove
Within-archive co-occurrence only.

TEST-NDE-LIFE-REVIEW-V1

Needs designQ-ONSET-GRAMMAR

Do key events appear in a repeatable order across different experience reports, and is that order different from fiction? This tests whether the sequence is real or only a keyword and storytelling effect.

Precise form: Do validated hallmark events have reproducible onset order across experiential corpora and fiction controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
OBERF/NDERF and fiction text exist
What's stopping us
context-aware extractor and gold labels
The boring explanation we must rule out first
broad keyword position masquerading as event order
How we'd check it honestly
Blinded event-identity precision gate before any order statistic
What would kill it
Kill if event precision fails or order matches fiction/source artifacts.
The most this could ever prove
Validated corpus order only; no universal experiential sequence.

TEST-ONSET-GRAMMAR-V1

Needs designQ-ONTOLOGY-SAFETY

Can broad labels for electrical, magnetic, and time-related experiences be split into clearer meanings without losing useful distinctions? Better labels would reduce false patterns caused by words with several meanings.

Precise form: Can broad EM/time labels be replaced without losing valid distinctions?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Old labels and source-grounded rows exist
What's stopping us
gold annotations
The boring explanation we must rule out first
polysemy and source-specific word meaning
How we'd check it honestly
Manual adjudication benchmark comparing broad and split ontologies
What would kill it
Reject any replacement that fails held-out meaning discrimination.
The most this could ever prove
Ontology-performance result only.

TEST-ONTOLOGY-SAFETY-V1

Needs designQ-SENSORY-REDUCTION

Does the same pattern appear across reduced-sensory experiments, hypnosis, dreams, and sleep paralysis after the collections are fairly matched? This would show whether it is specific to sensory reduction or common to many report types.

Precise form: Does a sensory-reduction pattern survive matched Ganzfeld, hypnosis, dream, and sleep-paralysis controls?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Several local corpora; protocol coverage uneven
What's stopping us
shared coding instrument
The boring explanation we must rule out first
report purpose, length, and vocabulary leakage
How we'd check it honestly
Matched human-validated labels with corpus-held-out contrasts
What would kill it
Kill if controls reproduce the same pattern after matching.
The most this could ever prove
Corpus separation only; no sensory-gating mechanism.

TEST-SENSORY-REDUCTION-V1

Needs designQ-THREAT-PRESENCE-INVARIANCE

Do labels for threat and a sensed presence mean and perform the same way across different report collections? If not, comparisons may reflect genre and ordinary fear words.

Precise form: Are threat/presence labels measurement-invariant across source families?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Candidate corpora and proposed benchmark exist
What's stopping us
gold set and label rubric
The boring explanation we must rule out first
ordinary fear, entity nouns, genre conventions
How we'd check it honestly
Blinded multi-source gold set with per-source precision/recall floors
What would kill it
Kill invariance if accuracy or meaning shifts materially by source.
The most this could ever prove
Measurement result only; no entity interpretation.

TEST-THREAT-PRESENCE-INVARIANCE-V1

Needs designQ-TIME-MEASUREMENT-INVARIANCE

Do labels for unusual experiences of time mean the same thing across different report collections? If they change with source or wording, any cross-source pattern is unreliable.

Precise form: What survives when the exploratory time measurement invariance pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-TIME-MEASUREMENT-INVARIANCE-V1

Needs designQ-UAP-CRYPTID-ECOLOGY

Does the apparent link between reports of unexplained objects and unusual creatures survive new examples checked by people and fair source controls? Shared vocabulary may create the pattern.

Precise form: What survives when the exploratory uap cryptid ecology pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-UAP-CRYPTID-ECOLOGY-V1

Needs designQ-UAP-EVIDENCE-CHANNEL

Can reports be separated by the kind of evidence they contain when tested on new examples checked by people? The first split may only reflect how each source asks people to report.

Precise form: What survives when the exploratory uap evidence channel pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-UAP-EVIDENCE-CHANNEL-V1

Needs designQ-UAP-REPORTING-ECOLOGY

Do different channels for reporting unexplained objects produce repeatable differences after source and writing style are controlled? This would show whether the pattern belongs to the reporting systems rather than the events.

Precise form: What survives when the exploratory uap reporting ecology pattern is tested on validated held-out labels?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Exploratory source artifacts exist
What's stopping us
validated labels and a frozen discriminant
The boring explanation we must rule out first
source vocabulary, report purpose, and unvalidated proxy labels
How we'd check it honestly
Frozen held-out label benchmark with source-matched controls
What would kill it
Kill the pattern if it fails held-out validated labels or source-matched controls.
The most this could ever prove
Exploratory corpus/measurement result only.

TEST-UAP-REPORTING-ECOLOGY-V1

Needs designQ-OBERF-TELLING-ORDER

Do people tell the main events in near-death reports in the same order in a second, independently collected archive? A match would support a reporting-order pattern, not the order in which the events were experienced.

Precise form: Does the coarse telling-order gradient replicate in the independently collected NDERF corpus?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Local NDERF exports (nderf_experiences.csv, nderf_reports.csv)
What's stopping us
second precision rater
The boring explanation we must rule out first
narrative prompts and questionnaire structure shaping telling order
How we'd check it honestly
Frozen extractor v2 + precision gate with a second rater + event-set-preserving permutation null
What would kill it
Kill if either endpoint fails p<=0.005 in both NDERF splits with gate-passing markers.
The most this could ever prove
Telling order in one additional corpus; still not experienced order.

TEST-OBERF-TELLING-ORDER-V1

Needs designQ-GANZFELD-DB-STRUCTURE

Does a second collection of telepathy experiments show the same overall excess of correct choices, no decline over time, and only a weak small-study pattern after duplicate studies are removed? This tests the compiled tables, not whether telepathy is real.

Precise form: Does the database-level structure (robust pooled excess, no decline, weak small-study signal) hold in the independent Tressoldi compilation?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Local MA_GanzfeldESP.xlsx (Tressoldi set)
What's stopping us
none
The boring explanation we must rule out first
overlapping study membership between compilations
How we'd check it honestly
Same frozen audit statistics plus a study-overlap dedup join between the two files
What would kill it
Kill consistency if conclusions flip after overlap dedup at frozen thresholds.
The most this could ever prove
Structure of compiled databases; upstream selection untested.

TEST-GANZFELD-DB-STRUCTURE-V1

Needs designQ-K7-N7-RAINBOW

Does an independent mathematician agree that the machine encoding and nine checked cases exactly answer the seven-vertex version of the problem? A correct machine proof can still answer the wrong formal question.

Precise form: Does human mathematical review confirm that the nine checked formulas exactly resolve Question 3.3 for n=7?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Public GitHub record plus retained CNF and DRAT files
What's stopping us
independent mathematician response
The boring explanation we must rule out first
a sound machine proof attached to an incomplete or incorrect translation of the original problem
How we'd check it honestly
Line-by-line expert review of the encoding, nine-case split, R3 semantics, small-weight closures, and symmetry transfer
What would kill it
Kill the theorem-level conclusion if any valid coloring pair is omitted or any transfer step is unsound.
The most this could ever prove
Machine-certified conditional result for n=7; no general odd-n claim.

TEST-K7-N7-RAINBOW-V1

Ready nowQ-JPL-CAD-CORRECTION

After later database updates, do the eight corrected close-approach records still keep the reported distance between their minimum and maximum values? This checks whether the repair stays in place.

Precise form: Do the eight corrected JPL close-approach records remain internally consistent after future data releases?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Public read-only checker and official JPL CAD API
What's stopping us
none
The boring explanation we must rule out first
a later orbit update or API change reintroducing a mismatch
How we'd check it honestly
Re-run the fixed eight-object invariant check and preserve a dated response receipt
What would kill it
Reopen the correction if any complete row again violates dist_min <= dist <= dist_max.
The most this could ever prove
Persistence of the correction in the eight reported objects only; no broader JPL reliability claim.

TEST-JPL-CAD-CORRECTION-V1

Needs dataQ-EVOBC-VARIANT-SCORING

How much do the published image-model scores change when each symbol variant is judged against a source-backed list of acceptable forms? Exact-copy scoring may mark valid variants wrong.

Precise form: How do the paper's trained-model scores change under a source-backed variant-aware evaluation?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Pinned EVOBC releases and frozen variant-sensitivity receipts; original split and predictions unavailable
What's stopping us
original evaluation split and per-image predictions
The boring explanation we must rule out first
the exact-copy lookup channel may not represent the trained models' errors
How we'd check it honestly
Hash-bind the original split and per-image predictions and rescore without changing the accepted-label table
What would kill it
Kill any trained-model correction claim if direct rescoring is negligible or reverses direction.
The most this could ever prove
Direct trained-model score change under a frozen source-backed relation table; no decipherment claim.

TEST-EVOBC-VARIANT-SCORING-V1

Needs dataQ-QUANTUM-SWITCH-REPRO

Does the reported quantum-device inequality stay above its target when measurement settings are randomized and results are checked in separate time blocks? This tests whether slow drift and fixed measurement order explain the result.

Precise form: Does the quantum-switch inequality remain above 1.75 when settings are randomized and analyzed in repeated time blocks?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Published aggregate YAML counts reproduce exactly; no repeated time blocks are public
What's stopping us
timestamped or repeated-block event data with setting order
The boring explanation we must rule out first
predefined setting order plus source or interferometer drift can bias aggregate conditional probabilities
How we'd check it honestly
Freeze randomized setting order and repeated blocks, then test the inequality within blocks and under order-preserving drift controls
What would kill it
Downgrade robustness if the block-aware lower bound reaches 1.75 or the violation tracks measurement order.
The most this could ever prove
Robustness of this apparatus and protocol to measured time-linked drift; no universal quantum-switch claim.

TEST-QUANTUM-SWITCH-REPRO-V1

Needs designQ-INEFF-FORM-HISTORY-2026-08-15

Did a change in the public near-death questionnaire make the link between ineffability and reported depth look stronger in later submissions? If so, an archive change could explain the trend without a population change.

Precise form: Does a documented NDERF questionnaire form change explain why the ineffability-depth association is stronger in later submissions?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
NDERF public site pages (lawful fetch); current export lacks form_version (empty in 4,965 of 4,966 rows)
What's stopping us
acquisition approval for public-page fetch
The boring explanation we must rule out first
form or prompt wording change, not a population change
How we'd check it honestly
Frozen form-history timeline from archived public pages; association strength recomputed within form periods; kill if strength moves exactly at a documented form boundary
What would kill it
association strength tracks documented form boundaries
The most this could ever prove
instrument-history explanation for one archive association

TEST-nde-archive-methods-V1

Needs designQ-RPREDICT-DEPTH-CORRELATES-2026-08-15

Do four tentative links between questionnaire answers and reported experience depth repeat in new records collected after August 15, 2026? Fresh records would show whether the first patterns survive a corrected analysis.

Precise form: Do the four exploratory depth-linked answers survive a repaired frozen rerun judged on NDERF records submitted after 2026-08-15?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
existing archive for design; fresh post-2026-08-15 submissions required for confirmation
What's stopping us
fresh NDERF submissions accumulating after 2026-08-15
The boring explanation we must rule out first
same-sitting response style and semantic overlap with the depth scale
How we'd check it honestly
formula-frozen covariates, decision-rule fixtures exercised pre-run, held-out fresh-record confirmation with Holm correction
What would kill it
any survivor loses significance or flips direction under the repaired pipeline or on fresh records
The most this could ever prove
associational archive structure only

TEST-rpredict-depth-correlates-V1

Needs designQ-RGEOM-CIRCUMSTANCE-2026-08-15

After reports are matched for overall depth, does the way a person nearly died change which features they report? This tests whether circumstance leaves a pattern or whether archive and form differences explain it.

Precise form: Does how a person nearly died leave any detectable fingerprint on the mix of features in their reported experience at matched depth?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
NDERF archive; a powered single-endpoint design is required because the first attempt could not detect its own minimum effect
What's stopping us
fresh submissions or an independent archive for confirmation
The boring explanation we must rule out first
hidden questionnaire form changes, small groups, and self-reported circumstance labels
How we'd check it honestly
single-endpoint powered test (transcendental subscale) with an outcome-blind power simulation frozen before any data is seen
What would kill it
the single-endpoint test fails at 0.05 on frozen fresh records or flips direction between eras
The most this could ever prove
structure of self-reported archive records only

TEST-rgeom-circumstance-fingerprint-V1

Ready nowQ-LAB-AME-CLOSURE-2026-08-22

Do all numbers derived in the published 2020 atomic-mass tables agree with the tables' own arithmetic at the shown precision? This checks the production of the files, not the underlying physics.

Precise form: Do the published AME2020 mass-table columns agree with their own arithmetic within printed precision?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
public AME2020 distribution files, hash-pinned
What's stopping us
The boring explanation we must rule out first
columns computed from one least-squares adjustment agree by construction except for production slips
How we'd check it honestly
exact-fraction recomputation of every derived column with planted-error, clean-fixture, and deletion controls
What would kill it
any row whose printed derived value is disjoint from recomputation at the one-unit envelope
The most this could ever prove
internal arithmetic structure of the published files only

TEST-LAB-AME-CLOSURE-2026-08-22-V1

Ready nowQ-LAB-AME-EDITIONS-2026-08-23

Why do two official versions of the 2020 atomic-mass table differ in the last printed digit for 16 values: production mistakes or different rounding at exact ties? Independent printings can show which explanation fits.

Precise form: Are the 16 last-digit disagreements between the two published AME2020 editions production slips or rounding ties?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
both editions plus the journal-printed table, a sibling file, and NUBASE2020, all hash-pinned before opening
What's stopping us
The boring explanation we must rule out first
near-boundary internal values printed under two rounding passes
How we'd check it honestly
frozen per-row wager over sealed independent printings with double-entered extraction
What would kill it
any independent printing siding against its own production line
The most this could ever prove
production and rounding structure of the 16 rows only

TEST-LAB-AME-EDITIONS-2026-08-23-V1

Needs designQ-LAB-ENSDF-CANDIDATES-2026-08-23

Will the nuclear-data maintainers confirm that 13 unexplained arithmetic conflicts are bookkeeping errors, or explain them as valid evaluation practice? Their row-by-row reply is needed before any public claim.

Precise form: Do the unexplained ENSDF arithmetic contradictions survive maintainer adjudication as real bookkeeping errors?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
13 candidate records with hashed evidence packets, negative-search receipts, and blind-arm consensus
What's stopping us
NNDC response to the 2026-08-23 private-first note
The boring explanation we must rule out first
evaluation practice the audit did not recognize or issues already known internally
How we'd check it honestly
the maintainers own record-by-record reply
What would kill it
every record comes back known or explained
The most this could ever prove
archive bookkeeping structure of the flagged records only

TEST-LAB-ENSDF-CANDIDATES-2026-08-23-V1

Needs designQ-FORMAL-SOURCE-MATCH

Big AI labs keep computer-readable copies of famous math problems. Do those copies still match the original problem databases they were copied from, or have some quietly gone stale?

Precise form: How many entries in public formal mathematics collections disagree with their current upstream sources?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
public repositories and OEIS
What's stopping us
source bindings for unresolved files
The boring explanation we must rule out first
ordinary update lag between a source database and its formal copy
How we'd check it honestly
source-match recomputation with pinned hashes
What would kill it
audited entries match their current sources
The most this could ever prove
per-entry dated correction candidates only

TEST-FORMAL-SOURCE-MATCH-V1

Needs designQ-PACKING-CANDIDATE-RECORDS-2026-08-31

Can the catalog maintainer repeat the checks for the 600, 700, and 800-circle files and confirm whether each belongs in the catalog? Until that reply, all three remain candidates.

Precise form: Will the Packomania maintainer reproduce and accept the three packing candidates?

1 linked test · last triaged 2026-07-23

The honest state of this test
What data we already have
Three fixed coordinate files, their saved check reports, and a public checker that tests every wall and pair.
What's stopping us
The catalog maintainer has not yet replied with a decision.
The boring explanation we must rule out first
The public catalog can be stale, a private stronger file can exist, or the maintainer can use a different packing convention.
How we'd check it honestly
The catalog maintainer reruns the three files and checks them against the current catalog and any earlier private files.
What would kill it
Any circle crosses a wall, any pair overlaps, or an earlier file supports an equal or larger radius.
The most this could ever prove
Three candidate records submitted to the catalog maintainer for verification; no accepted-record or best-possible-packing claim.

TEST-PACKING-CANDIDATE-RECORDS-2026-08-31-V1

Looking for what already ran? Results live on Findings. Investigations that ended without an answer — and what would revive them — live on Cold Cases.