Findings

Outcomes, with their limits attached

A finding is a checked outcome. A claim is the exact sentence that outcome supports. “Discovery” is only a badge after a fresh novelty review—not a separate pile of content.

Supported findingNOVELChecked 2026-08-31

Three circle-packing candidate records sent for independent verification

All 600 circles of the candidate record packing drawn at their exact certified positions inside the unit square
This picture is the actual record file: the 600-circle candidate packing, every circle at its certified position.

The result: Aletheia produced fixed coordinate files for 600, 700, and 800 equal circles in a square. A check that does not round the coordinates found that every circle stays inside the square and no two circles overlap. Each file supports a radius just above the value printed in the public Packomania catalog. These are candidate records, submitted to the catalog maintainer for verification. They are not accepted records and do not prove the best possible packings.

What it means: For each of three crowded square puzzles, the new file fits all the circles without overlap and is a little better than the public catalog entry. The catalog maintainer now has the files to check.

What it is NOT: Not three accepted records, because the catalog maintainer has not confirmed them. Not proof of the best possible packings, because one working layout cannot rule out a better one. Not universal proof of novelty, because private or unindexed stronger files may exist.

Next test that could settle it: The catalog maintainer reruns the three coordinate files and compares them with the current catalog and any earlier private files. Any failed wall or pair check, or any earlier equal or stronger file, removes that candidate.

Open question: Will the Packomania maintainer reproduce and accept the three packing candidates?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Three fixed coordinate files fit 600, 700, and 800 equal circles in a square at checked radii above the values printed in the public Packomania catalog. They are candidate records, submitted to the catalog maintainer for verification. They are not accepted records, do not prove the best possible packings, and do not rule out earlier private or unindexed results.

Novelty: Three fixed coordinate files pass full wall and pair checks at slightly larger radii and were sent to the catalog maintainer.

Question: Q-PACKING-CANDIDATE-RECORDS-2026-08-31

Evidence grade: G2_SUPPORTED

Evidence records: ART-085

Related: DeepMind's Formal Conjectures carried a stale A157225 statement; reported privately, confirmed and merged, Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors, The world's atomic mass table is arithmetically airtight

Supported findingCORRECTION NOT DISCOVERYChecked 2026-08-28

DeepMind's Formal Conjectures carried a stale A157225 statement; reported privately, confirmed and merged

The result: An audit of 16 computer-readable math problems in Google DeepMind's public Formal Conjectures collection found one whose statement had gone stale: OEIS A157225's source had recorded a published counterexample (716,993,899) that the formal copy did not reflect. Aletheia independently recomputed the counterexample, reported it privately, and a repository maintainer confirmed and merged the fix (PR #5114, 2026-08-24).

What it means: Google DeepMind keeps a public library of math problems written for computers. One entry was out of date: mathematicians had already partly settled it and the library still showed the old version. We checked the new number ourselves, told the team quietly, and they confirmed it and fixed the library within days. It is the second time an institution has corrected its records after one of our reports.

What it is NOT: Not a new math discovery, because a mathematician published the key number first and our role was catching the stale copy and verifying it; and not evidence the collection is broadly wrong, because of 16 entries audited only this one reached a supported mismatch and fourteen could not be judged either way.

Next test that could settle it: Resolve source bindings for the 14 unadjudicated files and rerun the frozen source-match audit across the collection.

Open question: How many entries in public formal mathematics collections disagree with their current upstream sources?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): One dated repository correction, confirmed by the maintainers. Not a new mathematical discovery, not a claim about the rest of the collection (14 of 16 files could not be adjudicated either way).

Novelty: caught that the formal benchmark copy had not absorbed the update, recomputed the number independently, and obtained maintainer confirmation and a merged fix

Question: Q-FORMAL-SOURCE-MATCH

Evidence grade: G3_EXTERNALLY_CONFIRMED

Evidence records: ART-042, ART-043

Related: Three circle-packing candidate records sent for independent verification, Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors, The world's atomic mass table is arithmetically airtight

Supported findingNOVELChecked 2026-08-23

Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors

The result: The two published editions of the AME2020 mass table disagree on the final digit of 16 atomic masses at equal printed precision. A frozen check against the journal-printed table, a sibling file, and NUBASE2020 showed every printing loyal to its own production line: all 16 close as legitimate rounding-boundary cases with no demonstrable slip.

What it means: What looked like sixteen possible typos in the mass table turned out to be numbers sitting exactly on a rounding boundary, printed honestly by two different production passes.

What it is NOT: Not proof the sixteen rows are error-free, because a slip matching every printing of its own family would look identical; it is proof no public printing can show one.

Next test that could settle it: None; the class is resolved and closed.

Open question: Are the 16 last-digit disagreements between the two published AME2020 editions production slips or rounding ties?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Production and rounding structure of the 16 published rows only; a slip that matches every printing of its own family would look identical, so the claim is that no public printing can demonstrate one.

Novelty: first documentation and resolution of the equal-precision edition disagreements

Question: Q-LAB-AME-EDITIONS-2026-08-23

Evidence grade: G1_BOUNDED_NULL

Evidence records: ART-082, ART-080

Related: The world's atomic mass table is arithmetically airtight

Supported findingNOVELChecked 2026-08-22

The world's atomic mass table is arithmetically airtight

The result: Every derived column of the AME2020 atomic mass tables (binding, beta-decay, separation, and reaction energies) was recomputed from the primary masses with exact arithmetic: 51,937 checks per rounding envelope and zero inconsistent rows, with the whole audit independently reproduced by a second system.

What it means: The reference book of nuclear masses agrees with its own arithmetic everywhere we checked, which is a verified clean bill for a table thousands of papers copy from.

What it is NOT: Not proof the underlying measurements are right, because columns computed from the same fit must agree unless production slips; only production consistency was certified.

Next test that could settle it: Point the proven verifier at the human-curated ENSDF gamma-level archive, where agreement is not entailed by one adjustment (done: see the companion cards).

Open question: Do the published AME2020 mass-table columns agree with their own arithmetic within printed precision?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Internal arithmetic structure of the published AME2020 distribution files only; agreement between columns computed from one adjustment certifies production consistency, not experimental truth.

Novelty: a third-party audit of the published output files arithmetic, including the rounded edition, which no source was found to have done

Question: Q-LAB-AME-CLOSURE-2026-08-22

Evidence grade: G2_SUPPORTED_WITHIN_ONE_DATABASE

Evidence records: ART-084, ART-079

Related: Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors

Supported findingCONFIRMATION PHENOMENON KNOWNChecked 2026-08-15

Deeper experiences are the ones people call impossible to put into words

The result: In 4,630 NDERF records, people answering YES to "was the experience difficult to express in words?" score higher on the standard depth scale than NO answerers (medians 16 vs 14, effect size 0.126, p=1.4e-13, confirmed independently twice). Present in both submission eras, stronger later. Cannot-describe language runs 17.6% in these narratives vs 0.77% in tinnitus patients describing their own hard-to-convey symptom (descriptive; survives length matching). This is a quantitative confirmation of a known feature of deep experiences, with receipts: preregistered rules, a sealed blind coding packet judged by a fresh isolated coder at 30/30 on every load-bearing category, and four adversarial review cycles, one of which caught and retracted an era-analysis error.

What it means: People are not sprinkling 'indescribable' around at random. The ones who say words fail are the ones whose experiences went deepest, and that pattern is strong enough to build on: it means the failure of language is itself a structured, studyable part of these experiences.

What it is NOT: Not proof the experiences are truly beyond language, because 'words cannot describe it' is also a phrase people learn from books and other tellers, and the questionnaire asks the question directly which nudges yes answers. Not a claim about verified persons, because the archive cannot confirm identities, so counts are distinct records.

Next test that could settle it: Recover NDERF questionnaire form history to test whether an instrument change explains the era strengthening

Open question: Does a documented NDERF questionnaire form change explain why the ineffability-depth association is stronger in later submissions?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Self-report answer vs self-report depth score, same sitting, one self-selected archive; records not verified persons; no mechanism, causality, or novelty claim.

Novelty: population-scale quantification of the structured archive answer against Greyson totals with sealed blind coding, independent recomputation, and submission-date era analysis

Question: Q-INEFF-FORM-HISTORY-2026-08-15

Evidence grade: G2_within_one_archive

Evidence records: ART-045, ART-046

Related: Out-of-body reports share a coarse telling order, Structure audit of the 1974-2018 ganzfeld study database, De-identified dream reports can be linked across time

Supported findingNOVELChecked 2026-08-04

Variant-aware scoring shifts an EVOBC lookup score

The result: Across 100 fixed 9:1 simulations on EVOBC, accepting Unicode-listed variant labels raised an exact-copy lookup score by a median 2.4386 percentage points. A stricter simplified, traditional, and shape-variant rule raised it by 1.5416 points. The pinned Chinese-label character tree matched all 229,170 English-tree records and recovered 9,365 of 219,004 exact-image families assigned across labels. The effect on the paper's trained models is unknown because the original split and per-image predictions remain unavailable.

What it means: A published AI benchmark for reading ancient Chinese characters scores models as wrong when they give a valid variant form of the right answer. Accepting documented variants shifts a lookup score by around two percentage points, enough to matter in model comparisons.

What it is NOT: Not a correction of the paper's published model scores, because the original test split and predictions are not public. Not a claim that every variant reading is legitimate, because only Unicode-documented relations were accepted.

Next test that could settle it: Obtain the original evaluation split and per-image predictions and rescore them against a documented source-backed relation table.

Open question: How do the paper's trained-model scores change under a source-backed variant-aware evaluation?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): EVOBC exact-copy score sensitivity under the frozen Unicode variant rules and the pinned Chinese-label character-tree reconciliation. No correction to the paper's trained-model scores and no claim that every alternate reading is source-validated.

Novelty: I supply the EVOBC-specific 229,170-image census, frozen exact-copy scoring sensitivity, and pinned Chinese-label reconciliation.

Question: Q-EVOBC-VARIANT-SCORING

Evidence grade: G2_SUPPORTED_DETERMINISTIC_LOOKUP

Evidence records: ART-040, ART-041

Related: Eight JPL close-approach records corrected, K7 rainbow stacking conditional machine proof

Supported findingREDISCOVERY NOT NOVELChecked 2026-08-04

Public quantum-switch counts reproduce the published inequality score

The result: Two separate implementations reconstructed the public July 2025 quantum-switch counts and obtained 1.8328528, matching the paper's rounded 1.8328 against a fixed-order bound of 1.75. The paper's stated uncertainty of 0.0045 is larger than every count-only model tested here. This checks the released arithmetic. It cannot close the experiment's acknowledged detection, setting-choice, spacetime, or time-linked drift loopholes.

What it means: The released data behind a headline quantum-causality experiment checks out: two independent reconstructions of the public counts land exactly on the published score, comfortably above the classical bound. The arithmetic holds.

What it is NOT: Not confirmation of the experiment's big claim, because the acknowledged loopholes (detection, setting choice, drift) live upstream of the released counts. Not a replication, because no new data was collected.

Next test that could settle it: Release or collect repeated time blocks with randomized settings so time-linked drift and setting-order effects can be measured directly.

Open question: Does the quantum-switch inequality remain above 1.75 when settings are randomized and analyzed in repeated time blocks?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Exact arithmetic of the pinned Zenodo v2 aggregate counts and bounded count-model sensitivity only. No loophole-free indefinite-causal-order, no-drift, mechanism, or novelty claim.

Novelty: Aletheia independently reconstructs every released term, audits multiple count models, and freezes failure tests without claiming a new physical effect.

Question: Q-QUANTUM-SWITCH-REPRO

Evidence grade: G3_INDEPENDENT_RELEASED_COUNT_REPRODUCTION

Evidence records: ART-051, ART-052

Related: Eight JPL close-approach records corrected

Supported findingNOT APPLICABLE OPERATIONAL CORRECTIONChecked 2026-07-28

Eight JPL close-approach records corrected

The result: Aletheia identified eight internally inconsistent NASA/JPL close-approach records. JPL confirmed an issue affecting pathological cases with close approaches to multiple bodies within a short time window and fixed it for 367943, saying the other cases should be fixed as well. Aletheia's live check confirmed all eight reported records now satisfy the distance-order invariant.

What it means: Eight records in NASA/JPL's asteroid close-approach database contradicted themselves, the error report was sent in, and JPL confirmed the software issue and corrected the data. An outside one-person check improved a flagship public dataset.

What it is NOT: Not a sign of broader JPL unreliability, because the issue hit a narrow class of unusual multi-body cases. Not an astronomical discovery, because the errors were bookkeeping inconsistencies, not new objects.

Next test that could settle it: Re-run the public live checker after future JPL data releases and investigate any recurrence before making a broader reliability claim.

Open question: Do the eight corrected JPL close-approach records remain internally consistent after future data releases?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Eight reported records were internally inconsistent; JPL confirmed the relevant issue class and one corrected example; all eight now pass the live invariant. No claim about orbit errors, astronomical anomalies, or unrelated JPL systems.

Novelty: Aletheia supplied the reproducible eight-record diagnosis and independently verified the post-fix state.

Question: Q-JPL-CAD-CORRECTION

Evidence grade: G3_PROVIDER_CONFIRMED_CORRECTION

Evidence records: ART-035, ART-029

Related: Variant-aware scoring shifts an EVOBC lookup score, K7 rainbow stacking conditional machine proof

Supported findingNOVEL VERIFIEDChecked 2026-07-27

K7 rainbow stacking conditional machine proof

The result: All nine frozen machine cases for the seven-vertex rainbow-stacking problem are closed as unsatisfiable. Six have direct checked certificates. Three have checked symmetry-reduced certificates plus a reviewed transfer argument. This is a conditional resolution pending human review of the encoding and mathematical reductions.

What it means: A seven-vertex math problem was closed by machine: all nine required cases are proven impossible, six directly and three through a checked symmetry argument. Pending one human review of the setup, the question is resolved.

What it is NOT: Not yet a finished theorem, because a human mathematician must still validate the encoding and the reduction steps. Not a general method claim, because the certificates cover exactly these nine formulas.

Next test that could settle it: Human mathematician review of the full reduction from Question 3.3 to the nine formulas.

Open question: Does human mathematical review confirm that the nine checked formulas exactly resolve Question 3.3 for n=7?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): The nine machine cases are closed; resolution of Question 3.3 at n=7 remains conditional on human validation of the encoding, split, R3 semantics, small-weight closures, and symmetry transfer.

Novelty: Aletheia closes the nine-case machine search for n=7 with checked certificates; the theorem-level translation remains under human review.

Question: Q-K7-N7-RAINBOW

Evidence grade: G3_MACHINE_CERTIFIED_CONDITIONAL

Evidence records: ART-030, ART-031

Related: Variant-aware scoring shifts an EVOBC lookup score, Eight JPL close-approach records corrected

Supported findingNOVELTY UNRESOLVEDChecked 2026-07-24

Out-of-body reports share a coarse telling order

The result: Validated markers in 2,192 OBERF narratives follow a coarse order in the telling: onset/body markers early, tunnel and light mid-narrative, felt presence after them, return-to-body last; p=1e-4 in both corpus halves under an event-set-preserving null. Early-cluster internal order is not stable, so the supported claim is a coarse gradient, not a step-by-step sequence.

What it means: People telling out-of-body stories follow a shared coarse order: how it started comes early, tunnel and light land mid-story, a felt presence after them, and the return to the body last. Two thousand narratives beat shuffled versions of themselves at odds of one in ten thousand, in both halves of the archive.

What it is NOT: Not the order events were experienced in, because this measures how stories are told. Not a fine-grained sequence, because the early markers swap order freely and only the coarse gradient is supported.

Next test that could settle it: Frozen replication on the independently collected NDERF corpus with a second precision rater.

Open question: Does the coarse telling-order gradient replicate in the independently collected NDERF corpus?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Telling order of validated markers in one corpus; not an experienced-event sequence; early-cluster order unstable.

Novelty: Preregistered, precision-gated marker measurement with permutation nulls and split replication on an OBE corpus; telling-order claim ceiling

Question: Q-OBERF-TELLING-ORDER

Evidence grade: G2_SUPPORTED_WITHIN_ONE_CORPUS

Evidence records: ART-048, ART-049

Related: Structure audit of the 1974-2018 ganzfeld study database, De-identified dream reports can be linked across time, The psychedelic brain-complexity rise is mostly a spectral artifact

Supported findingPARTLY KNOWNChecked 2026-07-24

Structure audit of the 1974-2018 ganzfeld study database

The result: In the public 128-study ganzfeld compilation: pooled excess +376.2 hits over the ~5,459-trial chance expectation (p<1e-5), robust to greedy removal of 46 of 128 studies; decline over time absent at the frozen 0.01 threshold; small-study association suggestive only (p=0.019-0.047). Structure of the compiled file; upstream selection untested; no psi conclusion either way.

What it means: The public database of 128 telepathy-style ganzfeld studies really does contain more hits than chance predicts, and the excess survives removing a third of the studies and shows no fade over time. Whatever explains it, the compiled numbers themselves are not fragile.

What it is NOT: Not evidence for telepathy, because the audit tested only the structure of the compiled table and upstream publication selection remains untested. Not a final word, because the same frozen audit on an independent compilation is the named next step.

Next test that could settle it: Same frozen audit on the independent Tressoldi MA_GanzfeldESP compilation.

Open question: Does the database-level structure (robust pooled excess, no decline, weak small-study signal) hold in the independent Tressoldi compilation?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Structure of one proponent-compiled database; upstream publication/selection effects untested; no psi conclusion.

Novelty: Frozen pre-registered thresholds, greedy fragility count, duplicate sensitivity, and independent exact-tail reproduction on the 2018 compilation

Question: Q-GANZFELD-DB-STRUCTURE

Evidence grade: G2_SUPPORTED_WITHIN_ONE_DATABASE

Evidence records: ART-047, ART-050

Related: Out-of-body reports share a coarse telling order, De-identified dream reports can be linked across time, The psychedelic brain-complexity rise is mostly a spectral artifact

Supported findingPARTLY KNOWNChecked 2026-07-23

De-identified dream reports can be linked across time

The result: Closed-set pseudonymous linkage within DreamBank after obvious content clues were removed.

What it means: Dream reports are individually distinctive enough that the same anonymous dreamer can be re-identified across time inside an archive even after obvious clues are removed. Anonymized dream text is not as anonymous as it looks.

What it is NOT: Not real-name identification, because the test only matched pseudonymous records to each other inside one closed archive. Not a privacy verdict on all dream data, because open-set testing with separated recorders is the named next step.

Next test that could settle it: Independent open-set cohort with recorder/transcriber separation.

Open question: Does dream-report linkage generalize to independent open-set contributors?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Closed-set pseudonymous archive linkage, not real-name identification.

Novelty: Aletheia directly measures closed-set pseudonymous linkage across 20 contributors after removing obvious names and content clues, with an explicit privacy framing.

Question: Q-DREAM-PRIVACY-LINKAGE

Evidence grade: G2_SUPPORTED_WITHIN_ONE_ARCHIVE

Evidence records: ART-038, ART-039

Related: Out-of-body reports share a coarse telling order, Structure audit of the 1974-2018 ganzfeld study database, The psychedelic brain-complexity rise is mostly a spectral artifact

Read the full dream-linkage investigation →
Supported findingNOVELTY UNRESOLVEDChecked 2026-06-10

The psychedelic brain-complexity rise is mostly a spectral artifact

The result: In two raw psychedelic-EEG datasets (ketamine and DMT), apparent increases in signal-diversity complexity (Lempel-Ziv) are largely explained by flattening of the 1/f aperiodic slope; after controlling for the aperiodic exponent, no separable complexity effect remains in these datasets.

What it means: In ketamine and DMT brain recordings, the celebrated rise in signal "complexity" mostly disappears once you account for a simple change in the overall shape of the frequency spectrum. The flagship complexity story largely reduces to a known, simpler measurement effect in these datasets.

What it is NOT: Not a claim that psychedelics do nothing to the brain, because the spectral change itself is real and large. Not universal, because two public datasets were tested and a person-level preregistered replication is the named next step.

Next test that could settle it: Replicate prospectively with person-level raw EEG and preregistered spectral controls.

Open question: Does the LZc effect remain absent after preregistered aperiodic control in new raw EEG?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Bounded corpus/measurement result; no mechanism or universality.

Novelty: Aletheia tests the specific 1/f aperiodic exponent as the confound in both ketamine and DMT raw datasets and reports no separable drug effect after that control.

Question: Q-LZC-APERIODIC

Evidence grade: G2_SUPPORTED_EQUIVALENT

Evidence records: ART-058, ART-044

Related: Out-of-body reports share a coarse telling order, Structure audit of the 1974-2018 ganzfeld study database, De-identified dream reports can be linked across time

Supported findingNOVELTY UNRESOLVEDChecked 2026-05-30

Hessdalen sighting days track observing conditions, not geomagnetics

The result: In the full-date Hessdalen event corpus, dated observations behave as an observation-ecology result rather than a physical explanation: event days are not elevated on same-day Kp against matched +/-7/14/21/28 day controls, while precipitation and wind-gust metrics are lower than controls. This claim does not explain what the Hessdalen lights are, does not claim weather causes the phenomenon, and does not refute geomagnetic hypotheses generally; it only says same-day Kp is not elevated in this bridge. Weather evidence is Open-Meteo reanalysis-grid data, not direct station instrument logs for every event.

What it means: On days people report the Hessdalen lights, the geomagnetic index is no higher than on matched nearby days, but the weather is drier and calmer. Sighting days look like good-viewing days, which says the reporting record tracks when people can see, before it tracks anything about the lights.

What it is NOT: Not an explanation of the Hessdalen lights, because the analysis only tested day-matching against indexes and reanalysis weather. Not a refutation of geomagnetic ideas generally, because it covers same-day Kp in this one dated corpus with grid weather, not station instruments.

Next test that could settle it: Replicate with direct station data and independent observation dates.

Open question: Does the matched weather/Kp result replicate on direct station data and held-out dates?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Bounded corpus/measurement result; no mechanism or universality.

Novelty: Aletheia applies matched +/-7/14/21/28-day controls to dated reports and separates same-day Kp non-elevation from lower reanalysis precipitation/gust observation ecology.

Question: Q-HESSDALEN-ECOLOGY

Evidence grade: G2_SUPPORTED_EQUIVALENT

Evidence records: ART-071, ART-070

Related: Out-of-body reports share a coarse telling order, Structure audit of the 1974-2018 ganzfeld study database, De-identified dream reports can be linked across time