Papers
Topics
Authors
Recent
Search
2000 character limit reached

Truly Subquadratic 3SUM and Truly Subcubic APSP via Triangles in Sparse Lopsided Graphs

Published 5 Oct 2026 in cs.DS and cs.CC | (2610.06783v1)

Abstract: We give the first polynomial improvements over the textbook algorithms for $3$SUM and All-Pairs Shortest Paths (APSP): we show how to deterministically solve $3$SUM on nn integers of polynomial size in O(n<sup>1.9992)O(n<sup>{1.9992}) time and APSP on directed nn-vertex graphs with polynomially bounded integer weights in O(n<sup>2.9995)O(n<sup>{2.9995}) time. This refutes the $3$SUM and APSP hypotheses. Using known reductions, we also refute the real-valued versions of the $3$SUM and APSP hypotheses, the Exact Triangle hypothesis, the Zero-Weight kk-Clique hypotheses, and the three rectangular hinted Online Matrix--Vector conjectures of van den Brand, Nanongkai, and Saranurak, and we give polynomial speedups for a variety of other problems. All of these results follow from a single new algorithm for thin matrix products. Let XX be an N×DN\times D integer matrix and YY a D×ND\times N integer matrix with D≤N<sup>1/18D\le N<sup>{1/18}, and let WW be any set of at most N<sup>2/</sup>DN<sup>2/\sqrt</sup> D positions. We compute the entries (XY)[I,J](XY)[I,J], (I,J)∈W(I,J)\in W, in O(N<sup>2/D<sup>0.063)O(N<sup>2/D<sup>{0.063}) operations, which is polynomially less than the time needed to write down XYXY or to compute N<sup>2/</sup>DN<sup>2/\sqrt</sup> D inner products one by one. We design this algorithm by modifying a variant of Coppersmith's rectangular matrix multiplication algorithm, built from a ten-multiplication identity of Schönhage, to perform only the operations needed for the entries in WW, and show that few operations are needed. Interpreted as a graph algorithm, this solves the All-Edges Sparse Triangle problem in truly subquadratic time on sparse lopsided tripartite graphs where two parts have nn vertices but one part has n<sup>εn<sup>{\varepsilon} vertices for $\varepsilon&lt;0.12$. By known reductions, Exact Triangle, and hence $3$SUM and APSP, reduce to this problem. We also give a data structure version that answers queries for single entries of XYXY, not known in advance.

Summary

  • The paper introduces a deterministic algorithm for computing sparse entries of a thin matrix product X * Y with D=O(N^ε), achieving O(N^2/log²D) complexity.
  • It enables truly subquadratic computation for All-Edges Sparse Triangle, leveraging Schönhage's ten-multiplication identity to prune redundant computations.
  • This breakthrough offers polynomial improvements for integer-based 3SUM, APSP, and Exact Triangle, with exponents O(n^1.9992) and O(n^2.998), respectively.

Problem setting and principal claims

The paper identifies a new algorithmic regime for sparse outputs of highly rectangular matrix products and uses it to obtain polynomial improvements for several canonical fine-grained problems. Its central result is a deterministic algorithm for computing a prescribed sparse set of entries of a thin product XYXY, where XX is N×DN \times D, YY is D×ND \times N, and DD is a small polynomial in NN. In the principal parameter setting, if N≥D18N \ge D^{18} and the wanted set WW contains at most N2/DN^2/\sqrt D positions, then the entries XX0 for XX1 can be computed in

XX2

operations. A stronger data-structure formulation preprocesses XX3 and XX4 in XX5 time and answers any individual entry query in XX6 time. More generally, the method applies whenever XX7 for every XX8, with a polynomial saving depending on the density of the queried positions and the desired query time (2610.06783).

The immediate graph interpretation is a lopsided All-Edges Sparse Triangle algorithm. In a tripartite graph with parts XX9 and N×DN \times D0 of size N×DN \times D1, a middle part of size at most N×DN \times D2, and a prescribed set N×DN \times D3 of query pairs, the algorithm counts or detects common neighbors in the middle part for all pairs in N×DN \times D4. In the regime N×DN \times D5 and N×DN \times D6, this is truly subquadratic. This is significant because the balanced All-Edges Sparse Triangle problem has long served as a central hardness source for N×DN \times D7SUM, APSP, Set Disjointness, and related problems, while the lopsided instances produced by several established reductions had not been exploited algorithmically.

Composing the new triangle algorithm with known reductions yields the headline bounds:

Problem Input restriction Running time
N×DN \times D8SUM Polynomially bounded integers N×DN \times D9
APSP Polynomially bounded integer weights YY0
YY1-product Polynomially bounded integer entries YY2
Exact Triangle Polynomially bounded integer weights YY3
Real-valued YY4SUM Comparisons, additions, subtractions YY5 expected
Real-valued APSP Comparisons, additions, subtractions YY6 expected

The exponents in the introduction are stated with slightly different rounded versions depending on whether the basic sparse-product theorem or its stronger data-structure refinement is used. The improvements are small in absolute exponent but polynomial rather than subpolynomial, thereby contradicting the standard YY7SUM, APSP, and Exact Triangle hypotheses in their stated integer-input models.

The sparse thin-product theorem

The technical core is a selective matrix-multiplication algorithm. It does not assume structure in the entries of YY8 and YY9, nor in the set D×ND \times N0 beyond its cardinality. This distinguishes the result from conventional fast matrix multiplication, which computes every entry of D×ND \times N1 and therefore incurs an D×ND \times N2 output cost.

The construction adapts the second, tensor-based algorithm of Coppersmith for rectangular matrix multiplication. Its algebraic foundation is a ten-multiplication identity due to Schönhage. The identity simultaneously represents:

  1. an outer product of two length-three vectors, requiring nine scalar products; and
  2. an inner product of two length-four vectors, requiring four scalar products.

The identity expresses these thirteen products through ten bilinear multiplications, at the cost of additional error terms. Crucially, its variables divide into “outer” variables associated with the nine-term product and “inner” variables associated with the four-term product. The error terms contain an outer output variable together with at least one inner input variable. This separation makes the error terms harmless when the recursion is restricted to the particular output strings used to encode thin matrix products.

Applying the identity recursively for D×ND \times N3 levels produces a recursion tree with D×ND \times N4 leaves. Each leaf corresponds to a product of two encoded input values. The recursion simultaneously computes many matrix products whose inner dimension is D×ND \times N5, where D×ND \times N6 is the number of levels at which the inner product is selected. For a fixed set D×ND \times N7 of D×ND \times N8 inner levels, the recursion represents a product of an D×ND \times N9 matrix and a DD0 matrix, with DD1. The DD2 possible choices of DD3 are packed into one recursive computation.

The authors then tile the target DD4 product into products of shape DD5. A tile contains a grid of block products, each assigned to a distinct inner-level set DD6. This arrangement is important because the input array for a tile depends on one row band of DD7 and one column band of DD8. The corresponding encodings can therefore be computed once per band and reused across many tiles.

The decisive modification is to separate encoding, recursive multiplication, and decoding. Encoding computes, for each input array, all leaf values DD9 or NN0. Once these values are available, the algorithm can prune the recursive decoding computation: a recursive branch is visited only if at least one wanted output depends on it. The resulting cost is controlled not by the number of wanted entries multiplied by the number of leaves per entry, but by the size of the union of the leaf sets contributing to all wanted entries.

Leaf sharing and the source of the polynomial saving

The leaf-counting argument is the most distinctive combinatorial component of the paper. An output string with exactly NN1 inner positions depends on NN2 leaves. Computing each output independently would therefore cost roughly NN3, while direct inner-product evaluation would cost NN4. Neither bound yields the desired improvement.

The authors classify a leaf by its order. A leaf has order NN5 if it chooses a non-special term rather than NN6 at exactly NN7 of the NN8 inner levels. The private leaf of an output chooses NN9 at all inner levels and has order zero. Leaves of higher order are shared by more output strings because replacing a N≥D18N \ge D^{18}0 by one of the nine other terms makes the leaf compatible with multiple choices of the inner-level set.

For a fixed output, the number of order-N≥D18N \ge D^{18}1 contributing leaves is

N≥D18N \ge D^{18}2

The total number of order-N≥D18N \ge D^{18}3 leaves in the recursion is

N≥D18N \ge D^{18}4

With N≥D18N \ge D^{18}5, the global number of leaves decays geometrically:

N≥D18N \ge D^{18}6

Thus, low-order leaves are few per output but potentially numerous globally, whereas high-order leaves are numerous per output but sparse globally and heavily shared. Splitting at order N≥D18N \ge D^{18}7 gives

N≥D18N \ge D^{18}8

where N≥D18N \ge D^{18}9|U|=M/\sqrt DWW0O(M/D{1/18}).</sup>Thealgorithmconsequentlyevaluatesfewerthanoneleafperoutputentryofthefullproduct,despitecomputingarbitraryintegermatrixproducts.</p><p>Thisisthemechanismbehindthecentralnumericalsaving.Itisnotaconsequenceofafasterboundonordinarymatrixmultiplication;rather,itexploitstheoverlappatternbetweentheleafsetsassociatedwithdifferentoutputpositions.Thepaper’sexplicittheoremthereforedemonstratesaformofoutput−sensitivebilinearcomputationthatisunavailablefromstandardrectangularmatrixmultiplicationboundsalone.</p><h2class=′paper−heading′id=′data−structures−and−parameter−trade−offs′>Datastructuresandparametertrade−offs</h2><p>Thepaperextendsthebatchalgorithmtoonlinequeriesforentriesof.</sup> The algorithm consequently evaluates fewer than one leaf per output entry of the full product, despite computing arbitrary integer matrix products.</p> <p>This is the mechanism behind the central numerical saving. It is not a consequence of a faster bound on ordinary matrix multiplication; rather, it exploits the overlap pattern between the leaf sets associated with different output positions. The paper’s explicit theorem therefore demonstrates a form of output-sensitive bilinear computation that is unavailable from standard rectangular matrix multiplication bounds alone.</p> <h2 class='paper-heading' id='data-structures-and-parameter-trade-offs'>Data structures and parameter trade-offs</h2> <p>The paper extends the batch algorithm to online queries for entries of W$1 whose positions are not known during preprocessing. This requires replacing query-dependent pruning with query-independent aggregation.

The preprocessing partitions the leaves into two classes. Leaves below a switching order $W$2 are read individually during a query. Leaves of order at least $W$3 are grouped into “boxes,” each a structured subcube of the recursion tree. The value of a box is the sum of the products at all leaves in that subcube. Boxes are shared among many output entries and are computed by dynamic programming: a box with one free position is obtained by summing ten boxes in which that position is fixed.

For an output entry, the high-order leaves are partitioned into exactly $W$4 boxes, while the number of boxes stored globally is controlled by the decay of $W$5. This yields a preprocessing/query trade-off. If $W$6 and $W$7, then the query exponent and preprocessing saving depend on

$W$8

and on the decay ratio $W$9H) is the binary entropy function. The thinness threshold approaches

$N^2/\sqrt D$0

as $N^2/\sqrt D1D≤N<sup>0.1204−o(1)).</sup></p><p>Oneexplicitsetting,1D\le N<sup>{0.1204-o(1)}).</sup></p> <p>One explicit setting, N^2/\sqrt D$2 and $N^2/\sqrt D$3, yields preprocessing time $N^2/\sqrt D$4 and query time $N^2/\sqrt D$5. The same setting gives the main lopsided triangle bound

$N^2/\sqrt D$6

For $N^2/\sqrt D$7, the query term is dominated by the preprocessing term. The result is therefore simultaneously faster than evaluating each queried inner product directly and faster than computing the full product by known rectangular matrix multiplication whenever the middle dimension is sufficiently thin.

Reductions to Exact Triangle, $N^2/\sqrt D$8SUM, and APSP

The graph algorithm is applied through a deterministic reduction from Exact Triangle. The reduction hashes edge-weight sums modulo a carefully selected prime $N^2/\sqrt D$9 of size $X$00. The prime is chosen deterministically by counting false positives for every candidate prime and selecting one minimizing that count.

For each residue class of an edge weight and each piece of the third vertex part, the reduction constructs a lopsided triangle instance whose middle vertices are pairs consisting of a graph vertex and a residue label. A query pair has a common middle neighbor exactly when the corresponding triangle has total weight zero modulo $X$01. The number of instances is at most $X$02, each with at most $X$03 query pairs. The false positives are subsequently scanned directly, and their total number is bounded by an averaging argument over the candidate primes.

Balancing the number of oracle instances against the witness-scanning cost gives an Exact Triangle algorithm in $X$04 time, with $X05O(n<sup>3−1/648log⁡<sup>2</sup></sup>n)05O(n<sup>{3-1/648}\log<sup>2</sup></sup> n) bound. The deterministic nature of the prime selection and the reduction is essential: it establishes a deterministic refutation of the integer Exact Triangle hypothesis rather than merely a randomized one.

The reductions from XX06SUM and APSP then transfer this improvement. For XX07SUM, the known reduction creates approximately XX08 Exact Triangle instances of size XX09, yielding an exponent improvement of approximately half the Exact Triangle saving. The paper obtains

XX10

For APSP and XX11-product, the standard reduction invokes Exact Triangle at size XX12, so only one third of the Exact Triangle exponent saving is retained. The resulting bound is

XX13

reported in the abstract as XX14. The same chain also gives truly subcubic algorithms for the broader APSP equivalence class, including Negative Triangle, Minimum Weight Cycle, Replacement Paths, Radius, Median, Tree Edit Distance, and several verification and witness-counting problems.

The reduction structure is algorithmically important. Many fine-grained reductions were developed to transfer conditional lower bounds and were not optimized for preserving exponent improvements. Here, the losses are explicit: the Exact Triangle reduction preserves only part of the lopsided-triangle saving, and the reductions from XX15SUM and APSP preserve additional fractions. The paper therefore demonstrates both the utility and the current inefficiency of the existing reduction network.

For real-valued inputs, integer hashing is unavailable. The paper instead composes the thin-product algorithm with the randomized reductions of Chan, Vassilevska Williams, and Xu, which use Fredman’s trick to transform comparisons between sums into comparisons between row-dependent and column-dependent differences. The required comparison counts are reduced to selected entries of a thin integer matrix product using a block decomposition associated with Matoušek’s dominance-product technique.

With d=n<sup>1/40),</sup>theresultingalgorithmsuseonlycomparisons,additions,andsubtractionsonrealnumbersandareLasVegas:</p><ul><li>reald=n<sup>{1/40}),</sup> the resulting algorithms use only comparisons, additions, and subtractions on real numbers and are Las Vegas:</p> <ul> <li>real X$16SUM in $X$17 expected time;

  • real Exact Triangle, $X$18-product, and APSP in $X$19 expected time.
  • The algorithms never return an incorrect answer; randomization affects only the running time. The paper also states high-probability bounds obtained by restarting after an excessive running time.

    Further consequences include polynomial improvements for weighted zero-, min-, and max-weight $X$20-Clique through the classical Nešetřil–Poljak reduction. For constant $X$21$X$22$X$23\varepsilon_T&gt;0). The same framework gives a truly subquadratic algorithm for 3XOR with logarithmic dimension and refutes the thin-regime versions of three hinted Online Matrix–Vector conjectures. In particular, when the hint dimension is $X$24 with $X25O(n<sup>2−0.063τ)25O(n<sup>{2-0.063\tau})X26O(n<sup>1+0.437τ)26O(n<sup>{1+0.437\tau}) time. These bounds are faster than both obvious strategies—precomputing the full product and evaluating a query by a direct matrix-vector product.

    The consequences for dynamic data-structure lower bounds are localized. The main hinted-OMv parameter regimes used for dynamic matrix inverse, determinant, matching, and related problems typically have much thicker matrices and are not refuted. What fails are the conjectural lower bounds at the thin end of their trade-offs, where the middle dimension is below the threshold supported by the new matrix algorithm.

    Limitations and open questions

    The result does not yield a faster algorithm for balanced All-Edges Sparse Triangle, whose conjectured XX27 complexity remains unaffected. Nor does it directly improve problems known only to be 3SUM-hard or APSP-hard in the reverse direction, such as several geometric problems, dynamic graph problems, and some distance-oracle problems. In particular, the paper does not resolve whether three collinear points among XX28 planar points can be found in truly subquadratic time.

    The method is algebraic, has very large hidden constants, and is not presented as practical. Its applicability depends critically on the specific sparsity and sharing properties of Schönhage’s ten-multiplication identity. The authors do not establish that comparable pruning is possible for Coppersmith–Winograd identities or other tensor decompositions. The threshold XX29 is therefore a limitation of the construction, not a lower bound on sparse thin matrix products.

    Several hypotheses remain untouched, including SETH, Orthogonal Vectors, unhinted OMv, XX30-SUM and XX31-XOR for XX32, and 3SUM-Indexing. The paper emphasizes that these problems lack the combination of reductions to sparse triangles, efficient decision-tree algorithms, and nondeterministic algorithms that characterizes XX33SUM, Exact Triangle, and APSP. Whether the same algebraic technique can address any of these problems is left open.

    The paper also reports that the discovery involved Claude, an AI model developed by Anthropic, while the authors supplied the mathematical analysis, simplification, parameter optimization, reductions, exposition, and formal verification described in the manuscript. The methodology is part of the paper’s provenance, but the mathematical claims stand or fall on the stated algorithms and proofs. The Lean formalization covers the principal integer-input theorems and selected corollaries; it does not eliminate the need to inspect the modeling assumptions, imported prior reductions, or the asymptotic interpretation of the computational model.

    Conclusion

    The paper’s main contribution is an output-sensitive algorithm for sparse entries of thin matrix products, obtained by pruning a recursive implementation of Schönhage’s identity and exploiting the sharing of high-order leaves across many outputs. This yields truly subquadratic lopsided sparse-triangle algorithms and, through established reductions, the first polynomial improvements over the textbook exponents for integer XX34SUM and APSP. The resulting refutations are confined to specific fine-grained hypotheses and parameter regimes, but they demonstrate that reductions traditionally used to propagate hardness can also propagate a single algebraic speedup across a broad problem class. The central unresolved issue is whether the same approach can extend beyond lopsided instances and the XX35 thin-product regime (2610.06783).

    Whiteboard

    Explain it Like I'm 14

    1. What is this paper about?

    This paper presents a new way to solve some very difficult computer science problems slightly faster than anyone could before.

    The two most famous problems improved are:

    • 3SUM: Given many numbers, are there three that add up to zero?
    • All-Pairs Shortest Paths (APSP): Given a network, what is the shortest route between every pair of locations?

    For a long time, researchers believed that the best possible algorithms needed about:

    • n2n^2 steps for 3SUM
    • n3n^3 steps for APSP

    The paper shows that this is not quite true. Its algorithms are a little faster:

    • 3SUM can be solved in about n1.9992n^{1.9992} time.
    • APSP can be solved in about n2.9995n^{2.9995} time.

    These improvements may look tiny, but in theoretical computer science, improving an exponent—even by a small amount—is a major result.

    2. What questions did the researchers ask?

    The researchers wanted to answer several related questions:

    1. Can 3SUM be solved in truly less than quadratic time? “Quadratic” means roughly n2n^2 work. The paper asks whether the exponent can be made smaller than 2 by a fixed amount.
    2. Can APSP be solved in truly less than cubic time? “Cubic” means roughly n3n^3 work. The paper asks whether the exponent can be reduced below 3.
    3. Can one new technique improve many difficult problems at once? Many problems are connected through mathematical translations called reductions. If one problem becomes easier, these connections may allow other problems to become easier too.
    4. Can the researchers compute only the answers they actually need, instead of calculating everything? This is the key idea behind their method.

    3. How did they approach the problem?

    A useful graph problem: finding triangles

    The paper focuses on a graph problem called All-Edges Sparse Triangle.

    Imagine a graph split into three groups of dots. The researchers want to know, for many pairs of dots, whether the two dots share a neighbor in the third group. Such a shared neighbor creates a triangle.

    A normal method might check every possible combination. That can take too long.

    The paper studies a special case where:

    • Two groups are large.
    • The third group is much smaller.
    • Only some pairs of dots need to be checked.

    This uneven shape is why the graph is called lopsided.

    The matrix version

    The same task can be written using matrices. A matrix is a rectangular table of numbers.

    Suppose we multiply:

    • an N×DN \times D matrix XX
    • by a D×ND \times N matrix YY

    Here, DD is much smaller than NN. The product XYXY normally has N2N^2 entries.

    Instead of calculating the whole product, the researchers calculate only a selected set of entries—called the wanted entries.

    This is like checking only certain squares in a huge multiplication table rather than filling in every square.

    Their main result is that, under certain conditions, these selected entries can be computed in

    O(N2D0.063)O\left(\frac{N^2}{D^{0.063}}\right)

    steps.

    The exact exponent is not the main point for a young reader. The important idea is that this is polynomially faster than N2N^2, while still answering many questions.

    A clever matrix multiplication technique

    The method is based on an older technique for multiplying rectangular matrices, developed by Don Coppersmith. The researchers changed that method so it avoids calculations that affect entries nobody asked for.

    A helpful analogy is this:

    Suppose a teacher has a giant answer sheet but only asks you to find the answers to a few questions. Instead of solving every question, you organize your work so that you calculate only the requested answers.

    The paper uses algebraic identities—special mathematical formulas—to share work between many calculations. This lets one operation help answer many questions at once.

    Reductions: turning one problem into another

    The researchers also use reductions. A reduction is a way to convert one problem into another.

    For example:

    • A 3SUM problem can be transformed into a triangle-finding problem.
    • An APSP problem can also be transformed into a related triangle problem.
    • The new triangle algorithm can then solve the transformed problem faster.

    It is similar to translating a question into another language, solving it there, and translating the answer back.

    4. What did they find?

    Faster algorithms for 3SUM and APSP

    The main numerical results are:

    Problem Previous basic time New time in the paper
    3SUM About n2n^2 About n1.9992n^{1.9992}
    APSP About n3n^3 About n2.9995n^{2.9995}
    Exact Triangle About n3n^3 About n2.9983n^{2.9983}

    The paper also gives slightly different results for real numbers, using randomized methods whose answers are always correct but whose running time is expected to be fast. These are called Las Vegas algorithms.

    Many other problems also improve

    Because many problems reduce to 3SUM, APSP, or Exact Triangle, the new technique improves algorithms for several other tasks, including:

    • shortest-path problems
    • tree edit distance
    • certain versions of set intersection
    • weighted triangle problems
    • some knapsack problems
    • certain clique problems
    • some matrix multiplication and convolution problems
    • selected online matrix-vector problems

    The paper does not improve every difficult problem. For example, the results do not directly refute the main conjectures about:

    • CNF-SAT
    • Orthogonal Vectors
    • ordinary Online Matrix-Vector multiplication
    • kk-SUM for k≥4k \geq 4

    This is important because it shows that the new technique has limits.

    Long-standing “hardness” beliefs are disproved

    Computer scientists had proposed fine-grained complexity hypotheses. These are beliefs that certain problems cannot be solved noticeably faster than their known algorithms.

    This paper disproves the 3SUM and APSP hypotheses for the stated types of inputs.

    That does not mean 3SUM or APSP are now easy. The improvement is very small, and the new algorithms may be impractical for normal-sized inputs. However, it proves that the old running times were not the final mathematical limit.

    The role of artificial intelligence

    The paper says that an AI model called Claude discovered the main algorithm. The authors then checked, explained, improved, and extended the idea.

    The paper also says that Claude helped verify the main results using Lean 4, a computer program that checks mathematical proofs very carefully.

    The authors remain responsible for the final paper and its correctness.

    5. Why are these results important?

    The biggest lesson is that reductions are useful in two directions.

    Before this work, researchers mainly used reductions to say:

    “If problem A is difficult, then problem B must also be difficult.”

    This paper shows another possibility:

    “If problem B becomes easier, then the reduction gives us a faster algorithm for problem A.”

    One new algorithm for a special triangle problem therefore improves many other problems connected to it.

    The paper also teaches researchers to look more carefully at sparse information. Standard matrix multiplication calculates every entry of the answer, even when only a few entries matter. This research shows that avoiding unnecessary entries can lead to real improvements.

    6. Simple conclusion

    This paper makes a small but important crack in some long-standing barriers in algorithm design. It shows that 3SUM and APSP can be solved slightly faster than previously believed, and that one clever algebraic technique can help many related problems.

    The algorithms are currently complicated and probably too slow for everyday use. Still, the discovery changes what researchers believe is possible. It may lead to:

    • faster practical algorithms in the future,
    • improved methods for sparse matrix calculations,
    • new ways to solve graph and optimization problems,
    • and a better understanding of which problems are genuinely difficult.

    In short, the paper shows that problems once thought to be stuck at n2n^2 or n3n^3 time can sometimes be improved by calculating only what is needed and by using hidden connections between different problems.

    Knowledge Gaps

    Knowledge gaps, limitations, and open questions

    The paper establishes polynomial improvements for several problems, but leaves the following issues unresolved:

    • The optimal exponent for thin matrix products is unknown. The main algorithm applies when D≤NεD\le N^{\varepsilon} with ε<0.1204\varepsilon<0.1204 and achieves a saving of DγD^{\gamma}, but it is unclear whether the method can handle larger values of DD, approach the known rectangular-multiplication threshold ε>0.321\varepsilon>0.321, or improve the savings beyond the stated constants.
    • The structural reason for the threshold ε<0.1204\varepsilon<0.1204 is not fully resolved. The limitation arises from the particular ten-multiplication Schönhage identity and its sparsity properties, but it remains open whether a different tensor identity or decomposition could extend the applicable range.
    • The algorithm does not improve the balanced All-Edges Sparse Triangle problem. The balanced regime, including instances with roughly equal tripartite parts and the conditional m4/3−o(1)m^{4/3-o(1)} barrier, remains unresolved.
    • The important regime with two parts of size nn and a third part of size n\sqrt n remains open. Even triangle detection, rather than the all-edges version, is not shown to admit an O(n2−δ)O(n^{2-\delta}) algorithm in this regime.
    • The paper does not determine the true complexity of All-Edges Sparse Triangle across general sparsity patterns. It leaves open whether a unified algorithm can interpolate between lopsided, balanced, and intermediate regimes.
    • The polynomial improvements are quantitatively small and may have very large hidden constants. The practical feasibility of the algorithms is not evaluated, and it is unknown whether the constructions can be simplified or implemented efficiently enough to outperform conventional algorithms at realistic input sizes.
    • No combinatorial analogue of the main algebraic algorithm is known. The results rely on integer matrix multiplication and algebraic identities; whether truly subquadratic or subcubic combinatorial algorithms exist for the affected problems remains open.
    • The paper does not establish lower bounds in the computational model used by the algorithm. Existing lower bounds for (min⁡,+)(\min,+)-product, path-comparison algorithms, and linear decision trees do not apply because the new algorithms transform weighted problems into unweighted counting or matrix-multiplication instances.
    • The exact status of the original APSP and $3$SUM hypotheses after these results is not replaced by a new robust hypothesis. The paper refutes the stated hypotheses but does not identify the correct fine-grained complexity or a new conjectured optimal exponent for APSP, $3$SUM, or Exact Triangle.
    • The results do not extend to CNF-SAT or Orthogonal Vectors. No reduction of these problems to the lopsided sparse-triangle setting is known, and the techniques do not address their apparently different dense color-triple or Boolean-vector structure.
    • The standard, unhindered Online Matrix–Vector conjecture remains unaffected. The data-structure results refute hinted variants only for thin hints; they do not provide an improvement for ordinary OMv.
    • The hinted OMv results cover only thin-hint parameter ranges. The conjectures remain open for the regimes relevant to several dynamic matrix inverse, dynamic distance, and dynamic graph lower bounds, particularly around hint dimensions near n0.5n^{0.5}.
    • Many 3SUM-hard and APSP-hard problems still lack faster algorithms. The paper emphasizes that one-way hardness reductions do not reverse into algorithms; examples include finding three collinear points, several geometric problems, and numerous dynamic graph problems.
    • The complexity of kk-SUM and kk-XOR for k≥4k\ge 4 remains unresolved. The techniques improve $3$SUM but do not yield algorithms near the usual n⌈k/2⌉n^{\lceil k/2\rceil} baselines for larger kk.
    • The lack of fine-grained self-reductions for kk-SUM and higher-arity kk-XOR remains a major obstacle. It is unknown whether suitable self-reductions exist and whether they would enable reductions or algorithms comparable to those for $3$SUM.
    • The $3$SUM-Indexing conjecture is not affected. The paper does not provide a faster data structure for the parameter regimes in which this conjecture is posed.
    • The real-valued algorithms are randomized and analyzed in expectation. The paper does not provide deterministic algorithms with comparable bounds for real-valued $3$SUM, APSP, or Exact Triangle.
    • The robustness of the real-valued reductions under stricter numerical models is unresolved. The results use comparisons, additions, and subtractions, but do not establish comparable guarantees under finite-precision arithmetic, numerical noise, or practical real-number representations.
    • The algorithms are restricted to polynomially bounded integer entries in their deterministic form. Their performance for larger integer magnitudes, arbitrary-precision inputs, or weights whose bit length is not logarithmic in the input size is not analyzed.
    • The data-structure trade-off is not known to be optimal. The preprocessing bound O(N2/D0.063)O(N^2/D^{0.063}) and query bound O(D0.437)O(D^{0.437}) are improvements over baseline methods, but no matching upper or lower bounds are given.
    • The data-structure model does not address dynamic updates. The results concern static preprocessing and queries; the cost of supporting updates to XX, YY, or the underlying set system remains open.
    • The paper does not clarify whether the sparse-output technique extends beyond counting products. It computes selected entries of ordinary integer matrix products, but analogous improvements for Boolean, min-plus, polynomial, or other semiring products are not established in general.
    • The consequences for graph girth approximation remain conditional on a further breakthrough. The paper notes that solving the n\sqrt n-middle-part triangle problem would affect girth approximation, but does not resolve either problem.
    • The improvements for derived problems depend on the quality and applicability of existing reductions. Problems known only to reduce from $3$SUM, APSP, or Exact Triangle do not automatically receive faster algorithms, leaving the algorithmic status of those problems unresolved.
    • The paper does not determine whether further improvements would propagate to currently unaffected conjectures. In particular, it remains unclear whether stronger sparse-triangle algorithms could eventually yield consequences for SETH, OV, unrestricted OMv, or higher-arity sum problems.
    • The broader limits of reduction-based fine-grained complexity are left open. The paper demonstrates that conditional hardness webs can become algorithmic transfer mechanisms, but does not characterize which classes of reductions are most likely to yield future simultaneous improvements across many problems.**

    Practical Applications

    Immediate Applications

    The paper’s strongest immediate value is as an algorithmic primitive for moderately thin, sparse matrix products and as a basis for improved exact algorithms. Although the asymptotic gains are currently small and the hidden constants may be large, the following uses are directly enabled by the reported results or by existing reductions.

    • Faster exact 3SUM and related numerical-search workloads — computational geometry and data analysis.
      • Potential workflow: use the new algorithm as a batch-search backend after normalizing inputs to bounded integers or to comparison/addition/subtraction operations.
      • Dependencies: inputs must satisfy the paper’s numerical assumptions; real-valued versions are randomized; the theoretical constants may make conventional sorting or hashing faster for practical input sizes.
      • Practical sectors: computational geometry, collision detection, constraint checking, symbolic data analysis, and scientific computing.
    • Improved exact APSP and (min⁡,+)(\min,+))-product computation — graph analytics and operations research.
      • Potential workflow: integrate the algorithm into exact graph-analysis libraries for dense or moderately dense directed networks, especially when all-pairs distances or min-plus products are required.
      • Applications: route-planning benchmarks, network reliability analysis, dependency graphs, scheduling, and dynamic-programming formulations based on min-plus algebra.
      • Dependencies: the improvements are polynomial but small; memory for storing n2n^2 distances remains unavoidable, and negative-cycle handling must still be performed or assumed absent.
    • Faster exact algorithms for APSP-equivalent graph problems — network and infrastructure analysis.
      • Potential products: graph-analysis toolkits offering exact diagnostic modules for transportation, communication, and dependency networks.
      • Dependencies: many reductions may introduce substantial overhead; improvements apply to the stated weight and graph regimes and do not automatically improve every sparse-graph implementation.
    • Improved directed unweighted APSP — software infrastructure and network services.
      • Potential workflow: use it in batch reachability-distance services, static dependency analysis, and large directed network audits.
      • Dependencies: the result is asymptotic and may not outperform highly optimized breadth-first-search implementations on sparse or moderate-sized graphs. The paper’s stated improvement is not a general breakthrough for online or dynamic shortest-path queries.
    • Thin hinted matrix–vector computation — databases, recommendation systems, and batched inference.
      • Potential workflow: preprocess a large collection of feature vectors or incidence matrices, then answer many restricted pairwise similarity, overlap, or count queries.
      • Examples: sparse set-overlap search, Boolean incidence queries, small-feature-dimension recommendation filters, and batched lookup services.
      • Dependencies: the dimension must be a sufficiently small power of NN; the queried entries must be sparse or handled individually; integer-size and word-RAM assumptions apply.
    • Offline set-disjointness and set-intersection counting — databases and information retrieval.
      • Potential tools: a batch set-intersection engine, graph-neighborhood overlap service, or incidence-matrix query index.
      • Applications: duplicate detection, document-term overlap, permission-set auditing, bipartite-neighborhood comparison, and join-like database workloads.
      • Dependencies: the universe size must be small relative to the number of sets, approximately within the paper’s thin regime. Large-universe, highly skewed, or online-adversarial workloads are not covered.
    • Triangle-through-edge counting in lopsided graphs — graph mining and anomaly detection.
      • Potential workflow: represent entities in two large groups and a small mediator/category group; construct biadjacency matrices; query selected pair counts.
      • Examples: shared suppliers between firms, shared tags between documents, common users between products, or common intermediaries in communication networks.
      • Dependencies: the graph must be structurally lopsided, with the small part roughly below the stated n0.12n^{0.12} threshold for the strongest general guarantee. Real-world graphs may not have this shape.
    • Improved exact knapsack and approximate subset-sum routines — logistics and resource planning.
      • Potential tools: capacity-indexed planning solvers, packing optimizers, and resource-allocation modules.
      • Dependencies: these are parameterized by capacity and may be useful only when the capacity is much smaller than the naive quadratic regime. The paper notes that some variants, including randomized $0/1$ knapsack results, are not uniformly deterministic.
    • Academic and policy use: revision of fine-grained complexity assumptions.
      • Actionable consequence: papers and software claims should no longer describe these hypotheses as credible unconditional barriers in the regimes refuted here.
      • Policy relevance: funding programs and algorithmic benchmark designers can treat these problems as active optimization targets rather than presumed-exponent frontiers.
      • Dependency: the result does not refute SETH, Orthogonal Vectors, unrestricted OMv, kk-SUM for k≥4k\ge4, or all 3SUM-hard problems. Lower-bound claims must be checked against the exact reduction direction and parameter regime.

    Long-Term Applications

    The following possibilities require further engineering, improved constants, extensions beyond the thin regime, or validation on realistic data.

    • Production-grade sparse matrix and graph-processing libraries.
      • Potential products: GPU/CPU libraries, compiler primitives, and graph-processing frameworks supporting “wanted-entry” matrix multiplication.
      • Required development: practical versions of the Schönhage/Coppersmith-based construction, memory-efficient layouts, parallelization, cache-aware implementations, and hardware-specific integer arithmetic.
      • Main obstacle: the paper explicitly notes enormous hidden constants and potentially impractical algebraic operations.
    • Large-scale database join and triangle-query engines.
      • Potential workflow: automatically detect lopsided join structure, convert relations into thin incidence matrices, compute only requested output pairs, and fall back to hash joins outside the valid regime.
      • Dependencies: data skew, update frequency, memory bandwidth, and the need to support non-integer attributes or approximate predicates.
    • Static recommendation, similarity, and bipartite-network analytics.
      • Potential tool: a batch “common-neighbor query accelerator” for selected pairs rather than all N2N^2 pairs.
      • Dependencies: privacy-preserving representations, rapidly changing data, and the requirement that the intermediary dimension remain sufficiently small. The paper provides no direct accuracy or privacy guarantees.
    • Exact optimization engines based on min-plus algebra.
      • Potential products: optimization solvers that select among classical dynamic programming, min-plus convolution, and the new algebraic routines based on parameter size.
      • Required development: extension to broader numeric ranges, floating-point robustness, negative and infinite values, parallel execution, and practical crossover analysis.
      • Dependency: algebraic speedups may be unsuitable when numerical stability or explainability is more important than asymptotic runtime.
    • Faster computational biology and structured sequence comparison.
      • Potential workflow: use the improved tree-edit-distance backend in RNA secondary-structure analysis or hierarchical document comparison.
      • Dependencies: real biological scoring schemes may use arbitrary or floating-point costs; practical performance depends on the reduction from tree edit distance and may not match the asymptotic graph algorithm directly.
    • Improved network planning and resilience analysis.
      • Potential applications: identify critical roads or communication links, evaluate alternate routes, and measure network centrality under failures.
      • Required development: adapt exact dense-graph algorithms to sparse, dynamic, geographically embedded, or streaming networks; provide approximation and incremental-update variants.
      • Dependency: the paper’s improvements are primarily for static batch computation, not continuously changing infrastructure.
    • Dynamic data structures with thin hints.
      • Potential products: systems that preprocess a narrow family of possible updates or queries and then answer the realized query rapidly.
      • Examples: dynamic reachability with restricted update domains, fast query serving for narrow feature spaces, and specialized attention computations with a small candidate dimension.
      • Dependencies: the benefits occur only in thin-hint regimes. The paper explicitly leaves the main dynamic-matrix-inverse trade-offs, unrestricted OMv, and many dynamic graph lower bounds unaffected.
    • Hardware and accelerator design for sparse-output algebra.
      • Potential direction: FPGA, ASIC, or GPU accelerators for thin matrix products, modular arithmetic, and sparse output routing.
      • Required development: map the ten-multiplication identity and its pruning strategy onto parallel hardware, control intermediate-expression growth, and compare against tensor-core multiplication.
      • Dependency: hardware usefulness depends on workloads with stable dimensions and predictable wanted-entry patterns.
    • New algorithms for broader parameter regimes.
      • Research targets: balanced All-Edges Sparse Triangle, sparse-output rectangular multiplication for DD near N0.3N^{0.3}, dynamic updates, and kk-SUM or kk-XOR for k≥4k\ge4.
      • Dependency: the current technique relies on sparsity properties of a particular algebraic identity and does not automatically generalize.
    • Algorithmic benchmarking and education.
      • Potential outputs: teaching modules on sparse-output matrix multiplication, benchmark instances for lopsided triangle counting, and research software comparing algebraic and combinatorial approaches.
      • Dependency: benchmarks should report constants, memory use, numerical restrictions, randomization, and crossover sizes; asymptotic improvements alone may otherwise be misleading for daily computational practice.

    Glossary

    • All-Pairs Shortest Paths (APSP): The problem of computing shortest-path distances between every pair of vertices in a weighted graph. “Given an edge-weighted graph on nn vertices with no negative cycles, compute the shortest-path distance between every pair of vertices.”
    • All-Edges Sparse Triangle: The problem of determining, for every edge in a sparse graph, whether that edge belongs to a triangle. “In graph terms, offline Set Disjointness is the All-Edges Sparse Triangle problem: given a graph with mm edges, decide for every edge whether it lies in a triangle.”
    • Algebraic complexity: The study of computational complexity using algebraic operations and structures. “This also opens a number of research directions in fine-grained complexity, algorithm design, and algebraic complexity”
    • Biadjacency matrix: A matrix representing adjacency relations between vertices in two parts of a bipartite graph. “if XX and YY are the two biadjacency matrices”
    • Boolean matrix multiplication (BMM): Matrix multiplication over the Boolean semiring, using logical OR and AND instead of arithmetic addition and multiplication. “The known algorithms for combinatorial BMM”
    • Coppersmith–Winograd identities: Algebraic identities used to derive fast matrix multiplication algorithms. “prior to this work, the authors had tried approaches like this using the Coppersmith--Winograd identities”
    • Co-nondeterministic algorithm: An algorithm that efficiently verifies certificates proving that an instance is a NO-instance. “A nondeterministic algorithm for a decision problem verifies a proof for YES, and a co-nondeterministic algorithm verifies a given proof for NO.”
    • Combinatorial algorithm: An algorithm whose main operations are combinatorial rather than algebraic, often excluding fast algebraic matrix multiplication. “The new algorithms are algebraic and potentially impractical in their current form”
    • Convolution-3SUM: A constrained form of 3SUM involving indexed sequence elements whose indices add. “Convolution-3SUM (do x0,…,xn−1x_0,\dots,x_{n-1} satisfy xi+xj=xi+jx_i+x_j=x_{i+j} for some i,ji,j?)”
    • Decision tree: A computational model in which a problem is solved through a sequence of branching tests on the input. “$3$SUM (and more generally kk-SUM) and also Exact Triangle have very efficient {\em linear decision trees}.”
    • Exact Triangle: The problem of determining whether a weighted tripartite graph contains a triangle whose edge weights sum to zero. “given a tripartite graph with integer edge weights, is there a triangle whose three weights sum to zero?”
    • Fine-grained complexity: The study of computational complexity at the level of precise asymptotic exponents. “Both theories relate problems via reductions. FGC focuses on improvements in the exponent”
    • Fine-grained reduction: A reduction that preserves sufficiently precise running-time exponents between problems. “a fine-grained reduction from problem AA to problem BB with respect to running times a(n)a(n) and b(n)b(n)”
    • Fredman’s trick: An algebraic rearrangement that converts comparisons involving sums into comparisons of differences, often useful for real-valued inputs. “uses only additions, subtractions, and comparisons (via Fredman's trick”
    • Girth approximation: The approximation of the length of the shortest cycle in a graph. “would have consequences for girth approximation in undirected unweighted graphs”
    • Hinted Online Matrix–Vector multiplication (OMv): An online matrix–vector problem in which partial information about an arriving vector is supplied in advance. “This allows us to refute the hinted Online Matrix--Vector (OMv) conjectures”
    • Inner product: The sum of coordinate-wise products of two vectors. “computing N2/DN^2/\sqrt D inner products one by one”
    • Las Vegas algorithm: A randomized algorithm that always returns a correct answer, with randomness affecting only its running time. “for real numbers, a Las Vegas algorithm with O(n1.998)O(n^{1.998}) expected time”
    • Linear decision tree: A decision tree whose tests are linear functions of the input values. “it is now known that for every k≥3k\geq 3, kk-SUM has $2k$-linear decision trees”
    • Lopsided graph: A graph whose parts have substantially different sizes. “a lopsided version of All-Edges Sparse Triangle, in which the graph is tripartite and one of the three parts is significantly smaller than the other two.”
    • Matrix multiplication exponent: The exponent governing the asymptotic time required for multiplying square matrices. “where 2≤ω<2.3722\le\omega<2.372 is the matrix multiplication exponent.”
    • Merlin–Arthur protocol: A randomized proof-verification framework in which a prover supplies a certificate that an algorithm checks probabilistically. “the Merlin--Arthur world where randomization is allowed.”
    • Min-plus product: Matrix multiplication in which multiplication is replaced by addition and addition is replaced by minimum. “the (min⁡,+)(\min,+)-product A⋆BA\star B of two n×nn\times n matrices”
    • Nondeterministic algorithm: An algorithm that verifies a certificate for a YES-instance rather than finding a solution deterministically. “A nondeterministic algorithm for a decision problem verifies a proof for YES”
    • Orthogonal Vectors (OV): The problem of determining whether two Boolean vectors have no coordinate in which both contain a one. “Given nn Boolean vectors in d=ω(log⁡n)d=\omega(\log n) dimensions, are two of them orthogonal?”
    • Polynomial method: An algorithmic technique that represents combinatorial conditions using low-degree polynomials. “to design the fastest fine-grained algorithms for a number of problems using the polynomial method”
    • Rectangular matrix multiplication: Matrix multiplication involving matrices with substantially different dimensions. “the exponent of multiplying an n×nμn\times n^\mu by an nμ×nn^\mu\times n matrix.”
    • Reduction: A transformation that converts instances of one computational problem into instances of another while preserving relevant properties. “a reduction from AA to BB”
    • Semiring: An algebraic structure supporting addition-like and multiplication-like operations without requiring additive inverses. “over the (min⁡,+)(\min,+) semiring”
    • SETH: The Strong Exponential Time Hypothesis, which conjectures that CNF-SAT cannot be solved in substantially less than 2n2^n time. “the ``Strong Exponential Time Hypothesis,'' SETH”
    • Sparse matrix product: A matrix product in which only a selected subset of output entries is computed. “computing a prescribed sparse set of entries of a dense product of thin matrices.”
    • Sparse triangle: A triangle-finding problem in a graph with relatively few edges or a sparse set of relevant outputs. “All-Edges Sparse Triangle can be solved in O(m3/2)O(m^{3/2}) time just by listing all triangles”
    • Straight-line program: A sequence of arithmetic operations with no branching, used as a restricted computational model. “over the (min⁡,+)(\min,+) semiring, straight-line programs need n3n^3 operations”
    • Thin matrix: A matrix with one dimension that is a small power of the other, particularly an N×DN\times D matrix with small DD. “We call a product of an N×DN\times D matrix by a D×ND\times N matrix thin when DD is a small power of NN”
    • Truly subquadratic: Running in time O(n2−ε)O(n^{2-\varepsilon}) for some constant ε>0\varepsilon>0. “meaning in O(n3−ε)O(n^{3-\varepsilon}) time for some constant ε>0\varepsilon>0 (truly subquadratic is defined similarly).”
    • Truly subcubic: Running in time O(n3−ε)O(n^{3-\varepsilon}) for some constant ε>0\varepsilon>0. “all of the following problems are solvable in truly subcubic time”
    • Word RAM: A computational model in which memory words hold a logarithmic number of bits and basic word operations take constant time. “in the word RAM model of computation with O(log⁡n)O(\log n)-bit words”
    • Zero-Weight kk-Clique: The problem of finding a kk-vertex clique whose edge weights sum to zero. “Zero-Weight kk-Clique folding”
    • 3-linear degeneracy testing: Testing whether three input values satisfy a fixed linear relation. “all nontrivial variants of 3-Linear Degeneracy Testing”
    • 3SUM hypothesis: The conjecture that no truly subquadratic algorithm exists for 3SUM. “the ``$3$SUM hypothesis''”
    • 3XOR: A problem asking whether three vectors over F2\mathbb F_2 XOR to the zero vector. “Given three lists of nn vectors in F2d\mathbb F_2^{d}”

    Tweets

    Sign up for free to view the 18 tweets with 2704 likes about this paper.

    HackerNews

    1. Subquadratic 3SUM and Subcubic APSP (93 points, 36 comments) 
    2. Subquadratic 3SUM and Subcubic APSP (32 points, 3 comments)