FastQC: γ-Quasi-Clique Enumeration
- FastQC is a branch-and-bound algorithm that efficiently enumerates maximal γ-quasi-cliques in undirected graphs through advanced pruning and branching rules.
- It employs size–disconnection pruning and symmetry-exploiting candidate refinement to limit the search space and ensure the QC maximality property.
- The divide-and-conquer variant, DCFastQC, accelerates processing on large or dense graphs, achieving up to 100× speedup over previous approaches.
FastQC is a branch-and-bound (BB) algorithm designed for the efficient enumeration of all large maximal -quasi-cliques (MQCs) in undirected graphs. The method introduces co-designed pruning and branching rules to overcome time complexity barriers of prior approaches, complemented by a divide-and-conquer enhancement. FastQC enables the practical discovery of cohesive subgraph structures at scales and densities previously infeasible due to both theoretical worst-case bounds and empirical performance bottlenecks (Yu et al., 2023).
1. Problem Formulation: Maximal γ-Quasi-Clique Enumeration
Let be an undirected graph with vertices. Fix a density parameter and a size threshold .
- γ-Quasi-Clique (γ-QC): An induced subgraph () is a -quasi-clique if:
- is connected; and
- For every , 0.
Equivalently, defining the maximum number of disconnections in 1 as
2
we require 3, where 4.
- Maximal γ-Quasi-Clique (MQC): A 5-QC 6 is maximal if no 7 exists such that 8 is also a 9-QC.
- Enumeration Problem (MQCE): Output all MQCs in 0 of size at least 1.
Since QC maximality checking is NP-hard and the QC family is non-hereditary, enumeration is typically staged:
- Enumerate a superset 2 containing all MQCs.
- Filter 3 for maximality.
FastQC addresses the superset enumeration stage via an advanced BB framework.
2. Branch-and-Bound Framework in FastQC
FastQC explores the space of vertex subsets using a recursive BB search, where each node (branch) is represented by a triple 4:
- 5: current partial set (must-include vertices)
- 6: candidate set (potential additions to 7)
- 8: excluded set (must-not-include vertices)
Subsets 9 satisfying 0, 1, 2 are explored under each node.
The high-level recursive process FastQC-Rec(S, C, D) operates as follows:
- Apply pruning (described in Section 3).
- If candidate set 3 is empty, test 4 for being a 5-QC (output if 6).
- When an early-stop maximality criterion is satisfied, output 7 as a found MQC.
- Otherwise, select a pivot 8 and branch via the selected branching strategy (see Section 4).
3. Pruning Mechanisms: Size–Disconnection Space
FastQC introduces pruning rules in the "size–disconnection (SD)" space, efficiently restricting search to promising branches. For any 9:
- 0
- 1
For each node 2:
- Lower bound on QC size: 3
- Upper bound: 4, where 5
Branch 6 can only contain a 7-QC if:
8
Vertex refinement in 9 is performed by progressive pruning:
- (R1) Remove 0 if 1 — inclusion of 2 violates QC bounds.
- (R2) Remove 3 if 4 — no large QC could include 5.
These pruning operations are iterated, updating 6 and re-checking the pruning condition until no further removal is possible. If 7 at any step, the branch is safely pruned. This ensures the correctness and safety of the enumeration process, as only subsets 8 with 9 can form 0-QCs.
4. Branching Strategies: Sym-SE and Hybrid-SE
Once pruning and candidate refinement stall, FastQC branches the current search node using symmetry-exploiting rules.
- Symmetric SE branching (Sym-SE): Order candidates 1. Branch into 2 child nodes, each enforcing inclusion/exclusion decisions on 3. For increasing 4, 5 increases; if 6, all further siblings are pruned.
- Hybrid-SE branching: Triggered when a pivot 7 has 8 and 9. The "exclude 0" branch uses SE branching and a separate necessary condition for maximality; the "include 1" uses Sym-SE to prune non-QC branches—boosting pruning efficacy.
Pivot selection prefers 2 with maximal disconnections to maximize the utility of subsequent pruning.
5. Worst-Case Time Complexity
Let 3, with the bound 4 and typically 5. Define 6 as the maximal real root of 7. FastQC explores at most 8 recursion nodes, with 9 for all 0 (e.g., 1, 2, 3).
All pruning and branching rules are proven, via induction and recurrence analysis, to satisfy this exponential node growth, improving on the 4 worst-case bounds of prior BB approaches (Yu et al., 2023).
6. Divide-and-Conquer Enhancement: DCFastQC
To scale FastQC to massive or dense graphs, a divide-and-conquer variant ("DCFastQC", Editor's term) is employed:
- Compute a degeneracy ordering 5.
- For each 6, define 7 and form subgraph 8.
- Each MQC appears in exactly one 9 where 0 is minimal by degeneracy rank.
- Algorithm outline:
- Form 5;
- Iteratively prune via one-hop/two-hop pruning;
- Call `FastQC-Rec(S = {v_i}, C = V_i \setminus {v_i}, D = {v_1, \ldots, v_{i-1}})v \in H$6 and maximum degree $v \in H$7, each $v \in H$8 has $v \in H$9 vertices; aggregate runtime is $G = (V, E)$00, yielding substantial acceleration on sparse or locally bounded graphs.
7. Empirical Performance and Observations
Empirical evaluation on both real (12 network datasets, up to $G = (V, E)$01M vertices and $G = (V, E)$02M edges) and synthetic (Erdős–Rényi, up to $G = (V, E)$03 vertices, varied densities) graphs demonstrates:
- Against Quick+ (state-of-the-art BB), DCFastQC achieves up to $G = (V, E)$04 speedup on real-world networks (e.g., Enron, WordNet, Pokec), solving instances in seconds where Quick+ requires minutes or fails due to memory exhaustion.
- Speedup increases with $G = (V, E)$05 and $G = (V, E)$06 as the search space becomes more amenable to aggressive pruning.
- On dense Erdős–Rényi graphs, DCFastQC successfully handles $G = (V, E)$07–$G = (V, E)$08 vertices and edge-densities up to $G = (V, E)$09, where Quick+ fails beyond $G = (V, E)$10 vertices.
- DCFastQC and FastQC return identical sets of MQCs, but enumerate far fewer non-maximal QCs, significantly reducing the burden of the subsequent maximality filtering step and further improving overall enumeration efficiency.
These results establish FastQC (and DCFastQC) as among the fastest and most theoretically robust algorithms for maximal $G = (V, E)$11-quasi-clique enumeration in large-scale graphs (Yu et al., 2023).
References (1)