Every open question the lab has vetted, with its status. Each card expands to show the honest state of its test: what data exists, what is blocking it, the boring explanation that has to be ruled out first, and the most the test could ever prove even if everything goes right.
New here? What the labels mean
- Ready now —
- the data, the checks, and the kill rule all exist; this could run today.
- Needs data —
- the question is solid but a dataset is missing. If you have it, that is the whole blocker.
- Needs design —
- we do not yet have a test that could not fool itself.
- Needs a collaborator —
- requires an instrument, lab, or subject pool the lab does not have.
- “The most this could ever prove” —
- every test states its ceiling up front, so a narrow result can never quietly inflate into a big claim.
Needs dataQ-BIOLOGY-RECEPTOR
Do the 27 genes really form a biological link, or did name changes and mixed gene families create it? A false link could send later research in the wrong direction.
Precise form: Do receptor-family bridges survive exact 27-gene identity and provenance checks?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Existing 27-gene ledgers; expression and true receptor-family universes are absent
- What's stopping us
- expression matrix and receptor-family annotations
- The boring explanation we must rule out first
- gene-symbol aliases and mixed receptor families
- How we'd check it honestly
- Deterministic provenance join plus expression/family-matched nulls with planted controls
- What would kill it
- Kill any bridge whose members lack exact source-backed receptor-family identity.
- The most this could ever prove
- Dataset-specific provenance result; no biological mechanism.
TEST-BIOLOGY-RECEPTOR-V1
Needs dataQ-BODY-BOUNDARY
Do fainting, coma, anesthesia, sleep paralysis, near-death experiences, and extreme acceleration produce different patterns of feeling outside the body? Fair samples could show whether experiences that sound alike have different triggers.
Precise form: Do G-LOC, syncope, coma, anesthesia, sleep-paralysis, and NDE data preserve trigger-specific body-boundary patterns?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Some coded narratives; raw denominators incomplete
- What's stopping us
- raw denominators and independent labels
- The boring explanation we must rule out first
- keyword extraction, selected case series, unequal denominators
- How we'd check it honestly
- Blinded precision audit followed by diagnosis-matched proportion contrasts
- What would kill it
- Kill if validated labels or non-selected denominators remove the trigger contrast.
- The most this could ever prove
- Trigger/corpus contrast only; no shared mechanism.
TEST-BODY-BOUNDARY-V1
Needs designQ-COMMUNICATION-SPLIT
Can a computer separate reports with different kinds of communication when it is tested on new examples checked by people? The first pattern may only reflect how each source was written.
Precise form: What survives when the exploratory communication split pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-COMMUNICATION-SPLIT-V1
Needs dataQ-CORT-INTERMISSION
Do reports of an interval between lives form a repeatable pattern when every detail is traced to the original case record? This would show whether the proposed subtype is more than a set of memorable examples.
Precise form: Do intermission features survive a source-traceable case-level matrix?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Public CORT sources exist; case matrix incomplete
- What's stopping us
- case-level extraction
- The boring explanation we must rule out first
- secondary summaries and selected memorable cases
- How we'd check it honestly
- Double-entry source verification with blind feature coding
- What would kill it
- Kill if source-verifiable case rates do not reproduce the proposed subtype.
- The most this could ever prove
- Documented-case pattern only.
TEST-CORT-INTERMISSION-V1
Needs a collaboratorQ-CORT-PROSPECTIVE
Can claims about a previous life be recorded before anyone searches for a matching family, then match better than decoys in a blind test? This would test whether information leaks explain the apparent matches.
Precise form: Can prospective reincarnation-case documentation beat family and investigator information leakage?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Historical cases only
- What's stopping us
- participants and field collaborators
- The boring explanation we must rule out first
- retrospective memory, family cueing, investigator leakage
- How we'd check it honestly
- Prospective sealed statements, independent matching, and blinded scoring
- What would kill it
- Kill if blinded matches do not exceed preregistered decoys.
- The most this could ever prove
- Prospective case documentation only.
EXP-02
Needs a collaboratorQ-CRISIS-APPARITION-PROSPECTIVE
If people record unusual appearances when they happen, do those reports coincide with a death more often than chance and reporting opportunity allow? A modern test could avoid the memory and selection problems in old stories.
Precise form: Does a modern contemporaneous census show death-time coincidence beyond reporting opportunity?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Historical literature; no modern census
- What's stopping us
- participants, ethics, outcome linkage
- The boring explanation we must rule out first
- retrospective dating, denominator and reporting bias
- How we'd check it honestly
- Prospective timestamped reports linked blindly to outcome records
- What would kill it
- Kill if coincidence does not exceed the frozen opportunity-adjusted null.
- The most this could ever prove
- Prospective cohort association only.
EXP-03
Needs designQ-DISCOURSE-QUALITY
Can a computer separate the quality and structure of different reports when it is tested on new examples checked by people? The first pattern may only reflect writing style or source.
Precise form: What survives when the exploratory discourse quality pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-DISCOURSE-QUALITY-V1
Needs a collaboratorQ-DMT-NDE-BRIDGE
Do reports from the drug DMT and near-death experiences share anything specific after common words and general altered-state features are removed? This would test broad claims that the two experiences share one pattern.
Precise form: Do DMT and NDE reports share a specific signal beyond vocabulary and generic altered-state content?
4 linked tests · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Low-density/archive data tested; decisive paired data absent
- What's stopping us
- new high-density paired datasets
- The boring explanation we must rule out first
- item wording, vocabulary leakage, resolution mismatch
- How we'd check it honestly
- Held-out validated coding; physiology requires preregistered paired high-density data
- What would kill it
- Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
- The most this could ever prove
- Dataset-specific bridge or bounded null; no common mechanism.
EXP-BRIDGE-25-01
The honest state of this test
- What data we already have
- Low-density/archive data tested; decisive paired data absent
- What's stopping us
- new high-density paired datasets
- The boring explanation we must rule out first
- item wording, vocabulary leakage, resolution mismatch
- How we'd check it honestly
- Held-out validated coding; physiology requires preregistered paired high-density data
- What would kill it
- Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
- The most this could ever prove
- Dataset-specific bridge or bounded null; no common mechanism.
EXP-BRIDGE-25-02
The honest state of this test
- What data we already have
- Low-density/archive data tested; decisive paired data absent
- What's stopping us
- new high-density paired datasets
- The boring explanation we must rule out first
- item wording, vocabulary leakage, resolution mismatch
- How we'd check it honestly
- Held-out validated coding; physiology requires preregistered paired high-density data
- What would kill it
- Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
- The most this could ever prove
- Dataset-specific bridge or bounded null; no common mechanism.
EXP-BRIDGE-25-03
The honest state of this test
- What data we already have
- Low-density/archive data tested; decisive paired data absent
- What's stopping us
- new high-density paired datasets
- The boring explanation we must rule out first
- item wording, vocabulary leakage, resolution mismatch
- How we'd check it honestly
- Held-out validated coding; physiology requires preregistered paired high-density data
- What would kill it
- Kill if overlap vanishes under leakage controls or physiology fails preregistered targets.
- The most this could ever prove
- Dataset-specific bridge or bounded null; no common mechanism.
EXP-BRIDGE-25-04
Needs designQ-DREAM-EXPERIENCE-EEG
Does the brain signal linked to reporting a dream reflect dreaming itself, or only lighter sleep and greater alertness? Separating these would show whether the signal is specific.
Precise form: Is the dream-report EEG correlate more than arousal or sleep depth?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Archive data and exploratory result exist
- What's stopping us
- analysis design and adequate raw channels
- The boring explanation we must rule out first
- stage, time, absolute power, and arousal
- How we'd check it honestly
- Stage/time-matched replication with absolute power and source localization
- What would kill it
- Kill consciousness-specific framing if arousal/depth controls absorb the effect.
- The most this could ever prove
- Archive-local physiological correlate only.
TEST-DREAM-EXPERIENCE-EEG-V1
Needs dataQ-DREAM-PRIVACY-LINKAGE
Can dream reports be linked to the same writer in a new group when the possible author is not limited to a small known set? This matters because reliable linkage could create a privacy risk.
Precise form: Does dream-report linkage generalize to independent open-set contributors?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- DreamBank result only; external cohort absent
- What's stopping us
- independent consented cohort
- The boring explanation we must rule out first
- recurring content, recorder/transcriber, closed-set advantage
- How we'd check it honestly
- Open-set external cohort with contributor and recorder separation
- What would kill it
- Kill generalization if open-set linkage falls to the frozen null.
- The most this could ever prove
- Current finding remains one-archive closed-set linkage.
TEST-DREAM-PRIVACY-LINKAGE-V1
Needs designQ-EM-MEASUREMENT-INVARIANCE
Do labels for unusual electrical or magnetic events mean the same thing across different report collections? If not, comparisons between collections may be artifacts of wording.
Precise form: What survives when the exploratory em measurement invariance pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-EM-MEASUREMENT-INVARIANCE-V1
Needs designQ-ENVIRONMENTAL-SUBSTRATE
Are unusual trickster-like reports more common in wet, cave, and forest settings after blind review and fair source controls? This tests whether the pattern is environmental or only a feature of the stories.
Precise form: Does the wet/cave/forest lower-trickster pattern survive blinded row adjudication and corpus controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Existing 120-row stratified packet
- What's stopping us
- coding rubric and second coder
- The boring explanation we must rule out first
- genre words, setting mentions, report length, and source channel
- How we'd check it honestly
- Two blinded human coders plus held-out adjudication agreement and within-corpus estimates
- What would kill it
- Kill if adjudicated positives collapse to lexical artifacts or within-corpus effects vanish.
- The most this could ever prove
- One packet/corpus ecology result; no environmental causation.
TEST-ENVIRONMENTAL-SUBSTRATE-V1
Needs a collaboratorQ-GCP-PROSPECTIVE
During a future major event, does a worldwide network of random-number devices behave unusually after hardware faults and event selection are fixed in advance? A prospective test avoids choosing the event or analysis after seeing the data.
Precise form: Does a preregistered future event produce network behavior beyond hardware and selection controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Retrospective archive only
- What's stopping us
- future events and live network
- The boring explanation we must rule out first
- event selection, device drift, archive sensitivity choices
- How we'd check it honestly
- Prospective preregistration with hardware diagnostics and untouched endpoint
- What would kill it
- Kill if prospective endpoint stays inside its frozen null.
- The most this could ever prove
- Prospective network statistic only.
EXP-01
Needs dataQ-HESSDALEN-ECOLOGY
Do reported lights in Hessdalen still show the same links to weather and solar activity when direct station records and new dates are used? This would test the first result with better local data.
Precise form: Does the matched weather/Kp result replicate on direct station data and held-out dates?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Current dated corpus; direct station coverage incomplete
- What's stopping us
- station logs and independent dates
- The boring explanation we must rule out first
- reanalysis-grid error and observer opportunity
- How we'd check it honestly
- Held-out case-crossover using station logs and observer-density controls
- What would kill it
- Kill extension if Kp elevation appears or weather differences vanish.
- The most this could ever prove
- Current finding stays reanalysis-based observation ecology.
TEST-HESSDALEN-ECOLOGY-V1
Needs designQ-LIGHT-SEPARATION
Can a computer separate reports about unusual lights from other reports when it is tested on new examples checked by people? The first pattern may only reflect source vocabulary.
Precise form: What survives when the exploratory light separation pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-LIGHT-SEPARATION-V1
Needs dataQ-LZC-APERIODIC
After ordinary changes in the brain-wave spectrum are removed, does the signal-complexity measure still show no drug effect in new recordings? This would test whether the earlier null survives a cleaner measurement.
Precise form: Does the LZc effect remain absent after preregistered aperiodic control in new raw EEG?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Two retrospective datasets; prospective replication absent
- What's stopping us
- independent raw EEG
- The boring explanation we must rule out first
- spectral slope, power, state mixing, and small subjects
- How we'd check it honestly
- Person-level preregistered mixed model with spectral-surrogate comparison
- What would kill it
- Falsify the current claim if a powered drug effect survives slope control.
- The most this could ever prove
- Current claim remains two-dataset measurement result.
TEST-LZC-APERIODIC-V1
Needs designQ-NDE-DREAM-MEANING
Do near-death reports and dream reports differ in meaning and communication after writing style, report length, and repeated authors are controlled? This tests whether the contrast belongs to the experiences or only to the archives.
Precise form: Do NDE and DreamBank meaning/communication contrasts survive source, length, and ontology controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- NDE and DreamBank text exist
- What's stopping us
- validated refined ontology
- The boring explanation we must rule out first
- source style, length, repeated people, and ontology leakage
- How we'd check it honestly
- Blinded codebook labels with held-out source-family tests
- What would kill it
- Kill if the contrast vanishes under source/length controls.
- The most this could ever prove
- Two-archive semantic contrast only.
TEST-NDE-DREAM-MEANING-V1
Needs dataQ-NDE-LIFE-REVIEW
Can the reported link between life reviews and encounters with deceased relatives be reproduced from checked, row-level near-death reports? The original percentages cannot yet be traced to a validated table.
Precise form: Is the reported life-review/deceased-relative subtype reproducible from validated NDE rows?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Claim summary exists; validated row packet not found
- What's stopping us
- recover coded rows
- The boring explanation we must rule out first
- generated labels and denominator loss
- How we'd check it honestly
- Blind label validation plus exact denominator recomputation
- What would kill it
- Kill if 58.14% and 16.279% cannot be reproduced on validated labels.
- The most this could ever prove
- Within-archive co-occurrence only.
TEST-NDE-LIFE-REVIEW-V1
Needs designQ-ONSET-GRAMMAR
Do key events appear in a repeatable order across different experience reports, and is that order different from fiction? This tests whether the sequence is real or only a keyword and storytelling effect.
Precise form: Do validated hallmark events have reproducible onset order across experiential corpora and fiction controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- OBERF/NDERF and fiction text exist
- What's stopping us
- context-aware extractor and gold labels
- The boring explanation we must rule out first
- broad keyword position masquerading as event order
- How we'd check it honestly
- Blinded event-identity precision gate before any order statistic
- What would kill it
- Kill if event precision fails or order matches fiction/source artifacts.
- The most this could ever prove
- Validated corpus order only; no universal experiential sequence.
TEST-ONSET-GRAMMAR-V1
Needs designQ-ONTOLOGY-SAFETY
Can broad labels for electrical, magnetic, and time-related experiences be split into clearer meanings without losing useful distinctions? Better labels would reduce false patterns caused by words with several meanings.
Precise form: Can broad EM/time labels be replaced without losing valid distinctions?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Old labels and source-grounded rows exist
- What's stopping us
- gold annotations
- The boring explanation we must rule out first
- polysemy and source-specific word meaning
- How we'd check it honestly
- Manual adjudication benchmark comparing broad and split ontologies
- What would kill it
- Reject any replacement that fails held-out meaning discrimination.
- The most this could ever prove
- Ontology-performance result only.
TEST-ONTOLOGY-SAFETY-V1
Needs designQ-SENSORY-REDUCTION
Does the same pattern appear across reduced-sensory experiments, hypnosis, dreams, and sleep paralysis after the collections are fairly matched? This would show whether it is specific to sensory reduction or common to many report types.
Precise form: Does a sensory-reduction pattern survive matched Ganzfeld, hypnosis, dream, and sleep-paralysis controls?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Several local corpora; protocol coverage uneven
- What's stopping us
- shared coding instrument
- The boring explanation we must rule out first
- report purpose, length, and vocabulary leakage
- How we'd check it honestly
- Matched human-validated labels with corpus-held-out contrasts
- What would kill it
- Kill if controls reproduce the same pattern after matching.
- The most this could ever prove
- Corpus separation only; no sensory-gating mechanism.
TEST-SENSORY-REDUCTION-V1
Needs designQ-THREAT-PRESENCE-INVARIANCE
Do labels for threat and a sensed presence mean and perform the same way across different report collections? If not, comparisons may reflect genre and ordinary fear words.
Precise form: Are threat/presence labels measurement-invariant across source families?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Candidate corpora and proposed benchmark exist
- What's stopping us
- gold set and label rubric
- The boring explanation we must rule out first
- ordinary fear, entity nouns, genre conventions
- How we'd check it honestly
- Blinded multi-source gold set with per-source precision/recall floors
- What would kill it
- Kill invariance if accuracy or meaning shifts materially by source.
- The most this could ever prove
- Measurement result only; no entity interpretation.
TEST-THREAT-PRESENCE-INVARIANCE-V1
Needs designQ-TIME-MEASUREMENT-INVARIANCE
Do labels for unusual experiences of time mean the same thing across different report collections? If they change with source or wording, any cross-source pattern is unreliable.
Precise form: What survives when the exploratory time measurement invariance pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-TIME-MEASUREMENT-INVARIANCE-V1
Needs designQ-UAP-CRYPTID-ECOLOGY
Does the apparent link between reports of unexplained objects and unusual creatures survive new examples checked by people and fair source controls? Shared vocabulary may create the pattern.
Precise form: What survives when the exploratory uap cryptid ecology pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-UAP-CRYPTID-ECOLOGY-V1
Needs designQ-UAP-EVIDENCE-CHANNEL
Can reports be separated by the kind of evidence they contain when tested on new examples checked by people? The first split may only reflect how each source asks people to report.
Precise form: What survives when the exploratory uap evidence channel pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-UAP-EVIDENCE-CHANNEL-V1
Needs designQ-UAP-REPORTING-ECOLOGY
Do different channels for reporting unexplained objects produce repeatable differences after source and writing style are controlled? This would show whether the pattern belongs to the reporting systems rather than the events.
Precise form: What survives when the exploratory uap reporting ecology pattern is tested on validated held-out labels?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Exploratory source artifacts exist
- What's stopping us
- validated labels and a frozen discriminant
- The boring explanation we must rule out first
- source vocabulary, report purpose, and unvalidated proxy labels
- How we'd check it honestly
- Frozen held-out label benchmark with source-matched controls
- What would kill it
- Kill the pattern if it fails held-out validated labels or source-matched controls.
- The most this could ever prove
- Exploratory corpus/measurement result only.
TEST-UAP-REPORTING-ECOLOGY-V1
Needs designQ-OBERF-TELLING-ORDER
Do people tell the main events in near-death reports in the same order in a second, independently collected archive? A match would support a reporting-order pattern, not the order in which the events were experienced.
Precise form: Does the coarse telling-order gradient replicate in the independently collected NDERF corpus?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Local NDERF exports (nderf_experiences.csv, nderf_reports.csv)
- What's stopping us
- second precision rater
- The boring explanation we must rule out first
- narrative prompts and questionnaire structure shaping telling order
- How we'd check it honestly
- Frozen extractor v2 + precision gate with a second rater + event-set-preserving permutation null
- What would kill it
- Kill if either endpoint fails p<=0.005 in both NDERF splits with gate-passing markers.
- The most this could ever prove
- Telling order in one additional corpus; still not experienced order.
TEST-OBERF-TELLING-ORDER-V1
Needs designQ-GANZFELD-DB-STRUCTURE
Does a second collection of telepathy experiments show the same overall excess of correct choices, no decline over time, and only a weak small-study pattern after duplicate studies are removed? This tests the compiled tables, not whether telepathy is real.
Precise form: Does the database-level structure (robust pooled excess, no decline, weak small-study signal) hold in the independent Tressoldi compilation?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Local MA_GanzfeldESP.xlsx (Tressoldi set)
- What's stopping us
- none
- The boring explanation we must rule out first
- overlapping study membership between compilations
- How we'd check it honestly
- Same frozen audit statistics plus a study-overlap dedup join between the two files
- What would kill it
- Kill consistency if conclusions flip after overlap dedup at frozen thresholds.
- The most this could ever prove
- Structure of compiled databases; upstream selection untested.
TEST-GANZFELD-DB-STRUCTURE-V1
Needs designQ-K7-N7-RAINBOW
Does an independent mathematician agree that the machine encoding and nine checked cases exactly answer the seven-vertex version of the problem? A correct machine proof can still answer the wrong formal question.
Precise form: Does human mathematical review confirm that the nine checked formulas exactly resolve Question 3.3 for n=7?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Public GitHub record plus retained CNF and DRAT files
- What's stopping us
- independent mathematician response
- The boring explanation we must rule out first
- a sound machine proof attached to an incomplete or incorrect translation of the original problem
- How we'd check it honestly
- Line-by-line expert review of the encoding, nine-case split, R3 semantics, small-weight closures, and symmetry transfer
- What would kill it
- Kill the theorem-level conclusion if any valid coloring pair is omitted or any transfer step is unsound.
- The most this could ever prove
- Machine-certified conditional result for n=7; no general odd-n claim.
TEST-K7-N7-RAINBOW-V1
Ready nowQ-JPL-CAD-CORRECTION
After later database updates, do the eight corrected close-approach records still keep the reported distance between their minimum and maximum values? This checks whether the repair stays in place.
Precise form: Do the eight corrected JPL close-approach records remain internally consistent after future data releases?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Public read-only checker and official JPL CAD API
- What's stopping us
- none
- The boring explanation we must rule out first
- a later orbit update or API change reintroducing a mismatch
- How we'd check it honestly
- Re-run the fixed eight-object invariant check and preserve a dated response receipt
- What would kill it
- Reopen the correction if any complete row again violates dist_min <= dist <= dist_max.
- The most this could ever prove
- Persistence of the correction in the eight reported objects only; no broader JPL reliability claim.
TEST-JPL-CAD-CORRECTION-V1
Needs dataQ-EVOBC-VARIANT-SCORING
How much do the published image-model scores change when each symbol variant is judged against a source-backed list of acceptable forms? Exact-copy scoring may mark valid variants wrong.
Precise form: How do the paper's trained-model scores change under a source-backed variant-aware evaluation?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Pinned EVOBC releases and frozen variant-sensitivity receipts; original split and predictions unavailable
- What's stopping us
- original evaluation split and per-image predictions
- The boring explanation we must rule out first
- the exact-copy lookup channel may not represent the trained models' errors
- How we'd check it honestly
- Hash-bind the original split and per-image predictions and rescore without changing the accepted-label table
- What would kill it
- Kill any trained-model correction claim if direct rescoring is negligible or reverses direction.
- The most this could ever prove
- Direct trained-model score change under a frozen source-backed relation table; no decipherment claim.
TEST-EVOBC-VARIANT-SCORING-V1
Needs dataQ-QUANTUM-SWITCH-REPRO
Does the reported quantum-device inequality stay above its target when measurement settings are randomized and results are checked in separate time blocks? This tests whether slow drift and fixed measurement order explain the result.
Precise form: Does the quantum-switch inequality remain above 1.75 when settings are randomized and analyzed in repeated time blocks?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Published aggregate YAML counts reproduce exactly; no repeated time blocks are public
- What's stopping us
- timestamped or repeated-block event data with setting order
- The boring explanation we must rule out first
- predefined setting order plus source or interferometer drift can bias aggregate conditional probabilities
- How we'd check it honestly
- Freeze randomized setting order and repeated blocks, then test the inequality within blocks and under order-preserving drift controls
- What would kill it
- Downgrade robustness if the block-aware lower bound reaches 1.75 or the violation tracks measurement order.
- The most this could ever prove
- Robustness of this apparatus and protocol to measured time-linked drift; no universal quantum-switch claim.
TEST-QUANTUM-SWITCH-REPRO-V1
Needs designQ-INEFF-FORM-HISTORY-2026-08-15
Did a change in the public near-death questionnaire make the link between ineffability and reported depth look stronger in later submissions? If so, an archive change could explain the trend without a population change.
Precise form: Does a documented NDERF questionnaire form change explain why the ineffability-depth association is stronger in later submissions?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- NDERF public site pages (lawful fetch); current export lacks form_version (empty in 4,965 of 4,966 rows)
- What's stopping us
- acquisition approval for public-page fetch
- The boring explanation we must rule out first
- form or prompt wording change, not a population change
- How we'd check it honestly
- Frozen form-history timeline from archived public pages; association strength recomputed within form periods; kill if strength moves exactly at a documented form boundary
- What would kill it
- association strength tracks documented form boundaries
- The most this could ever prove
- instrument-history explanation for one archive association
TEST-nde-archive-methods-V1
Needs designQ-RPREDICT-DEPTH-CORRELATES-2026-08-15
Do four tentative links between questionnaire answers and reported experience depth repeat in new records collected after August 15, 2026? Fresh records would show whether the first patterns survive a corrected analysis.
Precise form: Do the four exploratory depth-linked answers survive a repaired frozen rerun judged on NDERF records submitted after 2026-08-15?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- existing archive for design; fresh post-2026-08-15 submissions required for confirmation
- What's stopping us
- fresh NDERF submissions accumulating after 2026-08-15
- The boring explanation we must rule out first
- same-sitting response style and semantic overlap with the depth scale
- How we'd check it honestly
- formula-frozen covariates, decision-rule fixtures exercised pre-run, held-out fresh-record confirmation with Holm correction
- What would kill it
- any survivor loses significance or flips direction under the repaired pipeline or on fresh records
- The most this could ever prove
- associational archive structure only
TEST-rpredict-depth-correlates-V1
Needs designQ-RGEOM-CIRCUMSTANCE-2026-08-15
After reports are matched for overall depth, does the way a person nearly died change which features they report? This tests whether circumstance leaves a pattern or whether archive and form differences explain it.
Precise form: Does how a person nearly died leave any detectable fingerprint on the mix of features in their reported experience at matched depth?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- NDERF archive; a powered single-endpoint design is required because the first attempt could not detect its own minimum effect
- What's stopping us
- fresh submissions or an independent archive for confirmation
- The boring explanation we must rule out first
- hidden questionnaire form changes, small groups, and self-reported circumstance labels
- How we'd check it honestly
- single-endpoint powered test (transcendental subscale) with an outcome-blind power simulation frozen before any data is seen
- What would kill it
- the single-endpoint test fails at 0.05 on frozen fresh records or flips direction between eras
- The most this could ever prove
- structure of self-reported archive records only
TEST-rgeom-circumstance-fingerprint-V1
Ready nowQ-LAB-AME-CLOSURE-2026-08-22
Do all numbers derived in the published 2020 atomic-mass tables agree with the tables' own arithmetic at the shown precision? This checks the production of the files, not the underlying physics.
Precise form: Do the published AME2020 mass-table columns agree with their own arithmetic within printed precision?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- public AME2020 distribution files, hash-pinned
- What's stopping us
- The boring explanation we must rule out first
- columns computed from one least-squares adjustment agree by construction except for production slips
- How we'd check it honestly
- exact-fraction recomputation of every derived column with planted-error, clean-fixture, and deletion controls
- What would kill it
- any row whose printed derived value is disjoint from recomputation at the one-unit envelope
- The most this could ever prove
- internal arithmetic structure of the published files only
TEST-LAB-AME-CLOSURE-2026-08-22-V1
Ready nowQ-LAB-AME-EDITIONS-2026-08-23
Why do two official versions of the 2020 atomic-mass table differ in the last printed digit for 16 values: production mistakes or different rounding at exact ties? Independent printings can show which explanation fits.
Precise form: Are the 16 last-digit disagreements between the two published AME2020 editions production slips or rounding ties?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- both editions plus the journal-printed table, a sibling file, and NUBASE2020, all hash-pinned before opening
- What's stopping us
- The boring explanation we must rule out first
- near-boundary internal values printed under two rounding passes
- How we'd check it honestly
- frozen per-row wager over sealed independent printings with double-entered extraction
- What would kill it
- any independent printing siding against its own production line
- The most this could ever prove
- production and rounding structure of the 16 rows only
TEST-LAB-AME-EDITIONS-2026-08-23-V1
Needs designQ-LAB-ENSDF-CANDIDATES-2026-08-23
Will the nuclear-data maintainers confirm that 13 unexplained arithmetic conflicts are bookkeeping errors, or explain them as valid evaluation practice? Their row-by-row reply is needed before any public claim.
Precise form: Do the unexplained ENSDF arithmetic contradictions survive maintainer adjudication as real bookkeeping errors?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- 13 candidate records with hashed evidence packets, negative-search receipts, and blind-arm consensus
- What's stopping us
- NNDC response to the 2026-08-23 private-first note
- The boring explanation we must rule out first
- evaluation practice the audit did not recognize or issues already known internally
- How we'd check it honestly
- the maintainers own record-by-record reply
- What would kill it
- every record comes back known or explained
- The most this could ever prove
- archive bookkeeping structure of the flagged records only
TEST-LAB-ENSDF-CANDIDATES-2026-08-23-V1
Needs designQ-FORMAL-SOURCE-MATCH
Big AI labs keep computer-readable copies of famous math problems. Do those copies still match the original problem databases they were copied from, or have some quietly gone stale?
Precise form: How many entries in public formal mathematics collections disagree with their current upstream sources?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- public repositories and OEIS
- What's stopping us
- source bindings for unresolved files
- The boring explanation we must rule out first
- ordinary update lag between a source database and its formal copy
- How we'd check it honestly
- source-match recomputation with pinned hashes
- What would kill it
- audited entries match their current sources
- The most this could ever prove
- per-entry dated correction candidates only
TEST-FORMAL-SOURCE-MATCH-V1
Needs designQ-PACKING-CANDIDATE-RECORDS-2026-08-31
Can the catalog maintainer repeat the checks for the 600, 700, and 800-circle files and confirm whether each belongs in the catalog? Until that reply, all three remain candidates.
Precise form: Will the Packomania maintainer reproduce and accept the three packing candidates?
1 linked test · last triaged 2026-07-23
The honest state of this test
- What data we already have
- Three fixed coordinate files, their saved check reports, and a public checker that tests every wall and pair.
- What's stopping us
- The catalog maintainer has not yet replied with a decision.
- The boring explanation we must rule out first
- The public catalog can be stale, a private stronger file can exist, or the maintainer can use a different packing convention.
- How we'd check it honestly
- The catalog maintainer reruns the three files and checks them against the current catalog and any earlier private files.
- What would kill it
- Any circle crosses a wall, any pair overlaps, or an earlier file supports an equal or larger radius.
- The most this could ever prove
- Three candidate records submitted to the catalog maintainer for verification; no accepted-record or best-possible-packing claim.
TEST-PACKING-CANDIDATE-RECORDS-2026-08-31-V1