a population object quietly becomes a finite estimator
we break it with one tree, one batch, or one negative sample. finite forests · club · rényicl
findings
homeall 23 proven breaks, each one dissected: what breaks, why the proof failed, what survives, and the repair that restores the theorem. sorted by how far along our verification is.
every break here survived our own adversarial checking — exact arithmetic, independent reconstruction, and formal certificates where the math allows them. a few may be independently known; we audit novelty before we claim priority.
complete packages
verified, write-up in progress
verified, not written up yet
small fixes
not rebuilt yet
these breaks aren't scattered typos. nearly all of them are one of five identity mutations — the paper starts proving one object and finishes claiming the theorem for another. we found the pattern, turned each mutation into a detector, and aimed it at the literature. that's why this list keeps growing.
we break it with one tree, one batch, or one negative sample. finite forests · club · rényicl
we freeze the width, isolate the subspace, and preserve the bad event. mine · adanorm · dimensional collapse
we compute both laws and collide them. matrix-rényi · dag-moe · deep graph infomax
we build inputs that satisfy the printed assumptions and nothing more. randomized sketches · transformer satisfiability
we find the exact boundary case the approximation can't survive. deepnet · clipped variance · fixup
written up, witnessed, independently checked. ready to publish.
what breaks: the proposed information quantity is negative for every positive rényi order except α = 1 — even when y is a deterministic bijection of x.
why it failed: tensor-product entropy intuition was imported into a hadamard-product gram construction. the proof checked endpoints and asserted the interior.
what survives: the matrix entropy remains a statistic. shannon order α = 1 retains real subadditivity. experiments do not become numerically nonexistent.
the fix: use α = 1 for information semantics, or describe non-shannon variants as empirical scores unless positivity is proved in the actual gram domain.
if it gets fixed: an entire descendant literature would need to separate useful scores from false mutual-information semantics.
what breaks: grf and orf define finite ensembles, then prove consistency and asymptotic normality for an ideal average over every tree.
why it failed: finite computational randomness disappeared inside the proof before the asymptotic limit was taken. at b = 1, the printed conclusions fail exactly.
what survives: complete-forest theory may survive, and sufficiently large practical forests may behave well. the headline theorem simply does not prove that.
the fix: state the result for complete forests or require the finite-complete discrepancy to be oₚ(σₙ); pointwise, bₙσₙ² → ∞ is a natural target.
if it gets fixed: tree count becomes an inferential parameter with a visible monte-carlo error term instead of an invisible engineering knob.
the error and its counterexample held up under independent checking. the standalone write-up isn't done yet.
what breaks: a five-node dag and a six-node dag collide for every deterministic aggregator, refuting proposition 3.1 / a.2.
why it failed: the proof cited computation injectivity as graph injectivity, despite the cited source explicitly warning that sharing structure is forgotten.
what survives: the architecture and experiments. theorem 3.2 may have a repairable conclusion; theorem 3.3 requires a separate audit.
the fix: state that the encoder represents anonymous computation or the unfolding tree, or add an explicit node-sharing identity channel.
if it gets fixed: dag models would finally declare whether they preserve structure, computation, or unfolding instead of treating them as synonyms.
each one has an exact counterexample and a fix on record.
what breaks: the published statistic is biased by exactly (n−1)/n and can fall below true mutual information at n = 2.
why it failed: diagonal pairs in the negative term come from the joint distribution, not the product of marginals.
what survives: the population club inequality.
the fix: delete the diagonal and divide by n(n−1).
if it gets fixed: population bounds and finite training objectives stop being advertised as the same object.
what breaks: the paper's explicit definition demands a deterministic cutoff; the proof supplies a sample-dependent random cutoff.
why it failed: a quantifier crossed the probability-one event.
what survives: standard almost-sure consistency may survive.
the fix: move the cutoff inside the probability-one event.
if it gets fixed: the theorem says what the uniform law actually proves instead of a stronger impossible statement.
what breaks: rank-one anisotropic rows satisfy the printed tail definition but double a least-squares cost and violate the theorem.
why it failed: the statement assumes centered tail control; the proof silently uses iid identity covariance.
what survives: the standard isotropic sub-gaussian sketch theorem.
the fix: require iid isotropic rows.
if it gets fixed: randomness labels stop substituting for the covariance assumptions doing the actual work.
what breaks: a positive non-affine quadratic family satisfies the constraints at every fixed hidden width h ≥ 3.
why it failed: the appendix sends h to infinity with a uniform bound absent from the theorem.
what survives: the chosen rule and experiments.
the fix: add the missing width-uniform premise or drop uniqueness.
if it gets fixed: the method becomes a design choice rather than a uniquely derived law.
what breaks: with arbitrary embeddings, satisfiability can encode halting and is undecidable — contradicting the accepted theorem.
why it failed: the proof claims a finite codomain makes an arbitrary embedding sequence eventually periodic. it does not.
what survives: decidability under an effectively finite embedding representation may survive.
the fix: restrict how embeddings are represented, expose only fixed-width positions, and count that representation in the input.
if it gets fixed: verification theorems must specify executable syntax instead of accepting an arbitrary semantic oracle.
what breaks: a uniform bit with one negative gives log 2, not the claimed skew-kl value log(8/5).
why it failed: the critic depends on freshly sampled negatives and a proof coefficient has the wrong sign.
what survives: a population or infinite-negative relationship may survive.
the fix: separate the finite sampled objective from the population identity.
if it gets fixed: finite-negative training stops inheriting identities proved for a different probability object.
what breaks: an invertible binary-symmetric channel refutes the theorem's promised label separation.
why it failed: the proof creates a collider and loses the required markov relation.
what survives: stronger variants with additional assumptions may survive.
the fix: state the missing channel or sufficiency conditions.
if it gets fixed: an information ordering no longer masquerades as a universal task-relevance ordering.
what breaks: the paper's own scalar proxy violates the printed global bound by roughly 233×.
why it failed: an uncontrolled local taylor approximation became a global inequality; β > 1 also excludes the deployed β < 1 regime.
what survives: the architecture and empirical results are not automatically refuted.
the fix: supply a valid uniform remainder bound or narrow the theorem to a local regime.
if it gets fixed: optimization guarantees would apply to the initialization actually used by the model.
what breaks: in one dimension, key 2 and query 3 produce exponent 18xₜxₛ where attention requires 6xₜxₛ.
why it failed: the identity composes the key projection twice.
what survives: the intended theorem appears locally repairable.
the fix: reverse the kernel orientation.
if it gets fixed: the formal identity matches the computation it claims to explain.
what breaks: an exact two-dimensional instance has a negative eigenvalue while its only singular value grows forever.
why it failed: the initial state can lie entirely in a growing positive eigenspace.
what survives: collapse results with an excitation assumption may survive.
the fix: control initialization and competing expanding subspaces.
if it gets fixed: spectral signs stop being treated as dynamics without checking which modes are occupied.
what breaks: an exact two-point distribution has variance 1/(2n) against the printed upper bound 3/(8n).
why it failed: the clipping-interval width was not squared.
what survives: the corrected range-based variance argument.
the fix: square the range term and audit descendants repeating it.
if it gets fixed: reported finite-sample confidence bounds stop being too small.
what breaks: the universal no-potential claim fails in one dimension, where every continuous field has a primitive.
why it failed: the theorem omitted a dimension restriction.
what survives: a d ≥ 2 result may survive.
the fix: add d ≥ 2 and recheck every remaining hypothesis.
if it gets fixed: a grand impossibility theorem becomes an honest multidimensional theorem.
real errors, just small ones — kept narrow on purpose.
what breaks: a three-node graph has zero local-global mutual information while row-shuffle discrimination remains possible.
why it failed: the implemented corruption law is not the product of marginals used in the proof.
what survives: the discrimination objective; the broad conceptual criticism has prior art.
the fix: describe the actual group-discrimination law or sample the required product distribution.
if it gets fixed: the exact certificate sharpens an existing criticism without stealing priority.
what breaks: softmax pooling cannot distinguish one identical element from arbitrarily many copies.
why it failed: normalization deletes global multiplicity; the proof also omits mandatory layer normalization.
what survives: broader activation classes and implementations with an independent cardinality channel.
the fix: make cardinality explicit and narrow the universality claim to the implemented architecture.
if it gets fixed: models stop claiming 1-wl counting power from a replication-invariant operator.
what breaks: a squaring nonlinearity and a 2d sign-changing diagonal map refute the lemma.
why it failed: noninjective nonlinearities erase the distinction the proof needs.
what survives: the intended logistic-sigmoid case.
the fix: require the relevant injectivity property.
if it gets fixed: the local lemma matches the nonlinearity actually used.
what breaks: nonzero query and key matrices can multiply to zero, leaving a linear one-lipschitz block.
why it failed: factor-wise nonzero was mistaken for a nondegenerate product.
what survives: a theorem with a stronger interaction condition.
the fix: state nondegeneracy on the composed bilinear form.
if it gets fixed: the theorem excludes the exact degenerate case that falsifies it.
what breaks: a two-logit example breaks the claim under the infinity norm.
why it failed: the statement says any norm while the proof is euclidean.
what survives: the euclidean argument.
the fix: use the corresponding dual norm or state the euclidean result.
if it gets fixed: the theorem stops changing geometry by adjective.
what breaks: the spectral range and self-loop stay probability are both printed incorrectly.
why it failed: two local calculations were generalized without checking the boundary cases.
what survives: the central diffusion construction.
the fix: correct the spectrum and transition probability statements.
if it gets fixed: the surrounding interpretation becomes exact without pretending the method collapsed.
these come from an older report. we don't claim them until we've rebuilt the proofs ourselves.
what breaks: the legacy report describes a disconnected-image gap and a stronger first-failure-cardinality theorem.
why it failed: a local affine identity was propagated through neighborhoods that need not be connected.
what survives: unknown until the proof is reconstructed.
the fix: rebuild the source witness, finite certificate, and novelty audit before claiming ownership.
if it gets fixed: if reconstructed, this may turn an unbounded sufficiency question into a finite lattice certificate.
what breaks: the legacy report describes a missing oscillation term and a two-generator countermodel.
why it failed: a value at a regularized optimizer replaced an unregularized supremum.
what survives: the empirical results reportedly survive.
the fix: reconstruct from the accepted source and restore the missing term.
if it gets fixed: a transferable detector template, not yet an owned paper correction.