Findings

Outcomes, with their limits attached

A finding is a checked outcome. A claim is the exact sentence that outcome supports. “Discovery” is only a badge after a fresh novelty review—not a separate pile of content.

Supported findingNOVELChecked 2026-08-31

Three circle-packing candidate records sent for independent verification

All 600 circles of the candidate record packing drawn at their exact certified positions inside the unit square
This picture is the actual record file: the 600-circle candidate packing, every circle at its certified position.

The result: Aletheia produced fixed coordinate files for 600, 700, and 800 equal circles in a square. A check that does not round the coordinates found that every circle stays inside the square and no two circles overlap. Each file supports a radius just above the value printed in the public Packomania catalog. These are candidate records, submitted to the catalog maintainer for verification. They are not accepted records and do not prove the best possible packings.

What it means: For each of three crowded square puzzles, the new file fits all the circles without overlap and is a little better than the public catalog entry. The catalog maintainer now has the files to check.

What it is NOT: Not three accepted records, because the catalog maintainer has not confirmed them. Not proof of the best possible packings, because one working layout cannot rule out a better one. Not universal proof of novelty, because private or unindexed stronger files may exist.

Next test that could settle it: The catalog maintainer reruns the three coordinate files and compares them with the current catalog and any earlier private files. Any failed wall or pair check, or any earlier equal or stronger file, removes that candidate.

Open question: Will the Packomania maintainer reproduce and accept the three packing candidates?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Three fixed coordinate files fit 600, 700, and 800 equal circles in a square at checked radii above the values printed in the public Packomania catalog. They are candidate records, submitted to the catalog maintainer for verification. They are not accepted records, do not prove the best possible packings, and do not rule out earlier private or unindexed results.

Novelty: Three fixed coordinate files pass full wall and pair checks at slightly larger radii and were sent to the catalog maintainer.

Question: Q-PACKING-CANDIDATE-RECORDS-2026-08-31

Evidence grade: G2_SUPPORTED

Evidence records: ART-085

Related: DeepMind's Formal Conjectures carried a stale A157225 statement; reported privately, confirmed and merged, Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors, Thirteen unexplained arithmetic contradictions in the nuclear structure archive, reported privately, pending maintainer review

Supported findingCORRECTION NOT DISCOVERYChecked 2026-08-28

DeepMind's Formal Conjectures carried a stale A157225 statement; reported privately, confirmed and merged

The result: An audit of 16 computer-readable math problems in Google DeepMind's public Formal Conjectures collection found one whose statement had gone stale: OEIS A157225's source had recorded a published counterexample (716,993,899) that the formal copy did not reflect. Aletheia independently recomputed the counterexample, reported it privately, and a repository maintainer confirmed and merged the fix (PR #5114, 2026-08-24).

What it means: Google DeepMind keeps a public library of math problems written for computers. One entry was out of date: mathematicians had already partly settled it and the library still showed the old version. We checked the new number ourselves, told the team quietly, and they confirmed it and fixed the library within days. It is the second time an institution has corrected its records after one of our reports.

What it is NOT: Not a new math discovery, because a mathematician published the key number first and our role was catching the stale copy and verifying it; and not evidence the collection is broadly wrong, because of 16 entries audited only this one reached a supported mismatch and fourteen could not be judged either way.

Next test that could settle it: Resolve source bindings for the 14 unadjudicated files and rerun the frozen source-match audit across the collection.

Open question: How many entries in public formal mathematics collections disagree with their current upstream sources?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): One dated repository correction, confirmed by the maintainers. Not a new mathematical discovery, not a claim about the rest of the collection (14 of 16 files could not be adjudicated either way).

Novelty: caught that the formal benchmark copy had not absorbed the update, recomputed the number independently, and obtained maintainer confirmation and a merged fix

Question: Q-FORMAL-SOURCE-MATCH

Evidence grade: G3_EXTERNALLY_CONFIRMED

Evidence records: ART-042, ART-043

Related: Three circle-packing candidate records sent for independent verification, Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors, Thirteen unexplained arithmetic contradictions in the nuclear structure archive, reported privately, pending maintainer review

Supported findingNOVELChecked 2026-08-23

Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors

The result: The two published editions of the AME2020 mass table disagree on the final digit of 16 atomic masses at equal printed precision. A frozen check against the journal-printed table, a sibling file, and NUBASE2020 showed every printing loyal to its own production line: all 16 close as legitimate rounding-boundary cases with no demonstrable slip.

What it means: What looked like sixteen possible typos in the mass table turned out to be numbers sitting exactly on a rounding boundary, printed honestly by two different production passes.

What it is NOT: Not proof the sixteen rows are error-free, because a slip matching every printing of its own family would look identical; it is proof no public printing can show one.

Next test that could settle it: None; the class is resolved and closed.

Open question: Are the 16 last-digit disagreements between the two published AME2020 editions production slips or rounding ties?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Production and rounding structure of the 16 published rows only; a slip that matches every printing of its own family would look identical, so the claim is that no public printing can demonstrate one.

Novelty: first documentation and resolution of the equal-precision edition disagreements

Question: Q-LAB-AME-EDITIONS-2026-08-23

Evidence grade: G1_BOUNDED_NULL

Evidence records: ART-082, ART-080

Related: Thirteen unexplained arithmetic contradictions in the nuclear structure archive, reported privately, pending maintainer review, The world's atomic mass table is arithmetically airtight

Exploratory outcomeNOVELTY UNRESOLVEDChecked 2026-08-23

Thirteen unexplained arithmetic contradictions in the nuclear structure archive, reported privately, pending maintainer review

The result: Of 299,105 placed gamma records in the ENSDF nuclear structure archive, 61 fail their own level-scheme arithmetic beyond ten times the stated uncertainties. The archive explains 48 of them through its own footnotes and source structure, and 13 have no documented explanation and no consistent placement in any sibling dataset, with a blind second classification agreeing on all 13. They were reported privately to the maintainers first and their identities are withheld here pending that review.

What it means: Thirteen records in the world reference archive of nuclear structure contradict their own arithmetic with no explanation anywhere in the archive, and the people who maintain it now have the list.

What it is NOT: Not thirteen confirmed errors, because the maintainers may recognize some as known issues or evaluation practice, and their answer is the real test.

Next test that could settle it: The maintainers reply adjudicates each record as known, explained, or new; identities and per-record receipts publish after their review or a courtesy window.

Open question: Do the unexplained ENSDF arithmetic contradictions survive maintainer adjudication as real bookkeeping errors?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Archive bookkeeping structure of the flagged records only; dual-arm exploratory candidates whose decisive adjudication is the maintainers reply; no record is called a confirmed error.

Novelty: receipted record-level contradictions with negative explanation searches, awaiting maintainer adjudication

Question: Q-LAB-ENSDF-CANDIDATES-2026-08-23

Evidence grade: G1_EXPLORATORY

Evidence records: ART-083, ART-081

Related: Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors, The world's atomic mass table is arithmetically airtight

Supported findingNOVELChecked 2026-08-22

The world's atomic mass table is arithmetically airtight

The result: Every derived column of the AME2020 atomic mass tables (binding, beta-decay, separation, and reaction energies) was recomputed from the primary masses with exact arithmetic: 51,937 checks per rounding envelope and zero inconsistent rows, with the whole audit independently reproduced by a second system.

What it means: The reference book of nuclear masses agrees with its own arithmetic everywhere we checked, which is a verified clean bill for a table thousands of papers copy from.

What it is NOT: Not proof the underlying measurements are right, because columns computed from the same fit must agree unless production slips; only production consistency was certified.

Next test that could settle it: Point the proven verifier at the human-curated ENSDF gamma-level archive, where agreement is not entailed by one adjustment (done: see the companion cards).

Open question: Do the published AME2020 mass-table columns agree with their own arithmetic within printed precision?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Internal arithmetic structure of the published AME2020 distribution files only; agreement between columns computed from one adjustment certifies production consistency, not experimental truth.

Novelty: a third-party audit of the published output files arithmetic, including the rounded edition, which no source was found to have done

Question: Q-LAB-AME-CLOSURE-2026-08-22

Evidence grade: G2_SUPPORTED_WITHIN_ONE_DATABASE

Evidence records: ART-084, ART-079

Related: Sixteen last-digit disagreements between the mass table editions are rounding ties, not errors, Thirteen unexplained arithmetic contradictions in the nuclear structure archive, reported privately, pending maintainer review

Supported findingNOVELChecked 2026-08-04

Variant-aware scoring shifts an EVOBC lookup score

The result: Across 100 fixed 9:1 simulations on EVOBC, accepting Unicode-listed variant labels raised an exact-copy lookup score by a median 2.4386 percentage points. A stricter simplified, traditional, and shape-variant rule raised it by 1.5416 points. The pinned Chinese-label character tree matched all 229,170 English-tree records and recovered 9,365 of 219,004 exact-image families assigned across labels. The effect on the paper's trained models is unknown because the original split and per-image predictions remain unavailable.

What it means: A published AI benchmark for reading ancient Chinese characters scores models as wrong when they give a valid variant form of the right answer. Accepting documented variants shifts a lookup score by around two percentage points, enough to matter in model comparisons.

What it is NOT: Not a correction of the paper's published model scores, because the original test split and predictions are not public. Not a claim that every variant reading is legitimate, because only Unicode-documented relations were accepted.

Next test that could settle it: Obtain the original evaluation split and per-image predictions and rescore them against a documented source-backed relation table.

Open question: How do the paper's trained-model scores change under a source-backed variant-aware evaluation?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): EVOBC exact-copy score sensitivity under the frozen Unicode variant rules and the pinned Chinese-label character-tree reconciliation. No correction to the paper's trained-model scores and no claim that every alternate reading is source-validated.

Novelty: I supply the EVOBC-specific 229,170-image census, frozen exact-copy scoring sensitivity, and pinned Chinese-label reconciliation.

Question: Q-EVOBC-VARIANT-SCORING

Evidence grade: G2_SUPPORTED_DETERMINISTIC_LOOKUP

Evidence records: ART-040, ART-041

Related: K7 rainbow stacking conditional machine proof, The electromagnetic keyword fails its own consistency test

Supported findingNOVEL VERIFIEDChecked 2026-07-27

K7 rainbow stacking conditional machine proof

The result: All nine frozen machine cases for the seven-vertex rainbow-stacking problem are closed as unsatisfiable. Six have direct checked certificates. Three have checked symmetry-reduced certificates plus a reviewed transfer argument. This is a conditional resolution pending human review of the encoding and mathematical reductions.

What it means: A seven-vertex math problem was closed by machine: all nine required cases are proven impossible, six directly and three through a checked symmetry argument. Pending one human review of the setup, the question is resolved.

What it is NOT: Not yet a finished theorem, because a human mathematician must still validate the encoding and the reduction steps. Not a general method claim, because the certificates cover exactly these nine formulas.

Next test that could settle it: Human mathematician review of the full reduction from Question 3.3 to the nine formulas.

Open question: Does human mathematical review confirm that the nine checked formulas exactly resolve Question 3.3 for n=7?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): The nine machine cases are closed; resolution of Question 3.3 at n=7 remains conditional on human validation of the encoding, split, R3 semantics, small-weight closures, and symmetry transfer.

Novelty: Aletheia closes the nine-case machine search for n=7 with checked certificates; the theorem-level translation remains under human review.

Question: Q-K7-N7-RAINBOW

Evidence grade: G3_MACHINE_CERTIFIED_CONDITIONAL

Evidence records: ART-030, ART-031

Related: Variant-aware scoring shifts an EVOBC lookup score, The electromagnetic keyword fails its own consistency test

Exploratory outcomeNovelty not reviewedChecked 2026-05-30

The electromagnetic keyword fails its own consistency test

The result: The broad electromagnetic proxy appears in unrelated corpora, suggesting measurement-invariance failure rather than a real cross-domain electromagnetic mechanism.

What it means: The word patterns used to detect "electromagnetic effects" fire in collections that have nothing to do with each other. That means the detector is measuring language habits, not a real cross-domain electromagnetic signal. A useful negative: it killed a tempting bridge.

What it is NOT: Not a statement that electromagnetic effects never occur in any report, because the failure is in the measuring stick. Not final, because a manual reading of the split meanings is still the named next step.

Next test that could settle it: Manually adjudicate split EM meanings across independent corpora.

Open question: What survives when the exploratory em measurement invariance pattern is tested on validated held-out labels?

Receipts (exact limits and evidence records)

Claim ceiling (the strongest sentence this evidence supports): Exploratory substrate only.

Question: Q-EM-MEASUREMENT-INVARIANCE

Evidence grade: G1_EXPLORATORY

Evidence records: ART-067

Related: Variant-aware scoring shifts an EVOBC lookup score, K7 rainbow stacking conditional machine proof