GMEA: Group-based Multi-objective Attack for 3DGS
- The paper introduces a universal black-box attack against 3DGS watermarking by formulating watermark removal as a multi-objective optimization problem that balances fidelity with effective watermark destruction.
- It leverages an indirect objective that minimizes feature map variance from convolutional networks to blind watermark detectors while preserving model quality.
- The group-based optimization strategy partitions large 3D models into independent clusters, reducing the search space and computational complexity during the attack.
Group-based Multi-objective Evolutionary Attack (GMEA) is a universal black-box attack framework for 3D Gaussian Splatting (3DGS) watermarking systems. It is designed to remove invisible digital watermarks from 3DGS models without any knowledge of the watermark’s content, embedding method, or detection process. The framework formulates watermark removal as a large-scale multi-objective optimization problem that balances watermark destruction against visual fidelity, introduces an indirect objective that blinds the watermark detector by minimizing the standard deviation of features extracted by a convolutional network, and uses a group-based optimization strategy to partition a 3DGS model into multiple independent sub-optimization problems (Zeng et al., 10 Aug 2025).
1. Definition and attack setting
GMEA targets the robustness of digital watermarking techniques used in 3DGS, including methods that embed either 1D bitstreams or 2D images for copyright protection. Its threat model is entirely black-box: the attacker does not assume access to model internals, gradients, decoder architecture, watermark content, or embedding details. Within this setting, GMEA seeks to produce an adversarially perturbed 3DGS model that disables watermark recovery while preserving the rendered scene’s visual quality (Zeng et al., 10 Aug 2025).
The framework is notable for being presented as the first universal black-box attack framework for this problem class. In the formulation described for GMEA, the attacked object is the watermarked 3DGS model , and the output is an attacked model obtainable through allowed perturbations such as pruning or color shift. The optimization therefore operates simultaneously over attack efficacy and distortion control, rather than treating watermark removal as a single scalar objective (Zeng et al., 10 Aug 2025).
A common misconception is that “group-based” here refers to group labels, grouped targets, or grouped semantic classes. In GMEA, the grouping mechanism is instead a structural decomposition of the 3DGS model itself: the Gaussian primitives are partitioned into spatial clusters, and each cluster is optimized independently before recombination. That design choice is central to making evolutionary search feasible in a search space with potentially millions of Gaussians (Zeng et al., 10 Aug 2025).
2. Multi-objective formulation and indirect detector blinding
GMEA is posed as a constrained multi-objective optimization problem:
where denotes the set of valid attacked models obtainable from the watermarked model by allowed perturbations (Zeng et al., 10 Aug 2025).
The first objective, , measures visual quality loss. In the reported formulation it is a weighted combination of loss and SSIM over multiple camera views. The second objective, , quantifies watermark destruction, but not by directly querying decoder accuracy. Because the watermarking system is black-box, GMEA uses a model-agnostic indirect objective: it minimizes the dispersion of convolutional feature maps so that the downstream decoder becomes uninformative. The decision variables are a mask vector indicating which Gaussian kernels to prune and a perturbation vector altering the RGB values of the remaining Gaussians (Zeng et al., 10 Aug 2025).
The dispersion term is defined channel-wise over a pre-trained convolutional feature extractor :
0
Using this quantity, the watermark-destruction objective for a sub-model averages standard deviation over channels and viewpoints:
1
The paper’s Appendix states that reducing the variance of the feature maps to zero compresses the information entropy of the feature distribution, making it maximally uninformative and thereby breaking the watermark decoding pipeline (Zeng et al., 10 Aug 2025).
This indirect formulation distinguishes GMEA from attacks that directly optimize classifier outputs or decoder predictions. The operative assumption is not knowledge of a specific watermark decoder, but the broader dependence of watermark extraction on structured, high-variance convolutional features. In that sense, GMEA attacks the feature substrate on which detector-specific logic is built (Zeng et al., 10 Aug 2025).
3. Group-based decomposition of the 3DGS search space
A direct evolutionary search over all Gaussian kernels is described as infeasible because 3DGS models can contain very large numbers of primitives. GMEA addresses this with a group-based optimization strategy. The positions of the Gaussians are partitioned by applying K-Means clustering to the 2 coordinates, producing 3 disjoint spatial clusters 4. Each cluster defines a sub-model 5, and each sub-model is optimized independently (Zeng et al., 10 Aug 2025).
The final attacked model is reconstructed as the union of the optimized sub-models:
6
This decomposition reduces the number of variables in each evolutionary run, facilitates parallelization, and produces a reported trade-off between runtime and memory usage: increasing the group count 7 increases memory but decreases time. The reported experiments state that the group-based strategy enables attacking large models in reasonable time, about 20 min for 8, with suitable GPU resources (Zeng et al., 10 Aug 2025).
The significance of the grouping mechanism is methodological rather than merely implementational. Earlier multi-objective adversarial work had already established that population-based search is well suited to black-box settings and can maintain diverse solutions along a Pareto front (Suzuki et al., 2019). GMEA adapts that general principle to the structural scale of 3DGS by making the optimization problem decomposable at the level of Gaussian subsets rather than attempting a monolithic search over the entire scene (Zeng et al., 10 Aug 2025).
4. Evolutionary search procedure
GMEA employs a variant of NSGA-II as its multi-objective evolutionary algorithm. In the described encoding, each individual combines a binary pruning mask with a vector of color perturbations for the Gaussians in a cluster. Variation is performed with Simulated Binary Crossover (SBX) and polynomial mutation with small step size. Candidate solutions are evaluated using 9 and 0 over multiple rendered views, after which non-dominated sorting identifies Pareto-optimal candidates and crowd-distance-based diversity maintenance preserves coverage of the trade-off surface (Zeng et al., 10 Aug 2025).
This workflow places GMEA within a broader line of black-box adversarial optimization that treats attack construction as a genuine multi-objective problem rather than a weighted single-objective surrogate. Earlier work on adversarial example generation under black-box conditions used EMO to return diverse non-dominated solutions, including robust examples and DCT-based perturbations for high-resolution images (Suzuki et al., 2019). Related black-box attack frameworks likewise used multi-objective evolutionary search to balance attack success against visual quality or robustness, for example in attention-guided attacks on high-resolution images (Wang et al., 2021), unrestricted attacks based on filter sequences (Baia et al., 2021), timing-lag-robust speech attacks (Ishida et al., 2020), and dual-domain backdoor trigger search with non-dominated sorting and preference-based selection (Liu et al., 2024).
What distinguishes GMEA within this lineage is the combination of three design choices: a 3DGS-specific representation of perturbations as pruning and color modification, a detector-blinding objective based on feature-map standard deviation rather than direct decoder loss, and a group-based decomposition strategy tailored to very large 3D scene representations (Zeng et al., 10 Aug 2025).
5. Empirical behavior and attack–fidelity trade-offs
Experiments reported for GMEA show that it effectively removes both 1D and 2D watermarks from mainstream 3DGS watermarking methods while maintaining high visual fidelity. For 1D watermarking, the Bit Accuracy Rate (BAR) for the decoded watermark drops from approximately 100% to around 65%, while the Watermark Uncertainty Score (WUS) and Information Destruction Score (IDS) approach values corresponding to random guessing. For 2D watermarking, the SSIM and PSNR between the decoded watermark and the original are reported to be drastically reduced after attack, indicating that the watermark becomes visually unrecognizable or destroyed (Zeng et al., 10 Aug 2025).
The same experiments report that scene quality remains high after attack. Rendered outputs often retain SSIM above 95%, PSNR in the 30–38 dB range, and low MSE/RMSE. The paper also describes a Pareto-front pattern in which some solutions maximize watermark destruction with minor artifacts, some maximize fidelity with minor attacks, and a balanced solution effectively destroys the watermark with nearly no visual side effects. This is consistent with the stated role of multi-objective search: instead of collapsing efficacy and fidelity into a single coefficient-tuned objective, the algorithm preserves multiple trade-off solutions from which an attacker can choose (Zeng et al., 10 Aug 2025).
The ablation evidence is also structurally important. When the quality objective is removed, watermark removal becomes even stronger, but severe visual artifacts appear. That result supports the claim that the visual-quality term is not ornamental; it is the mechanism that constrains the search away from degenerate but destructive solutions. A plausible implication is that GMEA’s operational value lies less in raw removal strength than in its ability to preserve a high-quality, usable 3D asset after watermark destruction (Zeng et al., 10 Aug 2025).
6. Relation to adjacent attack paradigms and open issues
GMEA belongs to the family of black-box, population-based, multi-objective attacks, but its application domain differs substantially from the image-, speech-, and backdoor-oriented settings explored in earlier work. EMO-based adversarial example generation showed that multi-objective search can produce diverse solutions, including robust adversarial examples and frequency-domain perturbations (Suzuki et al., 2019). LMOA used attention maps and large-scale multiobjective evolutionary optimization for high-resolution images (Wang et al., 2021). MES-VCSP used NSGA-II and neighborhood search to discover variable-length composite semantic perturbation sequences (Sun et al., 2023). LADDER framed backdoor trigger design as a multi-objective optimization problem over effectiveness, dual-domain stealthiness, and robustness, solved through MOEA with non-dominated sorting and preference-based selection (Liu et al., 2024). GMEA extends this general methodological tradition to 3DGS watermark removal by shifting the optimization target from classifier behavior to watermark-decoder vulnerability and by introducing group-based partitioning as a primary scalability device (Zeng et al., 10 Aug 2025).
The reported limitations are practical and architectural. Despite the group-based acceleration, the evolutionary process remains moderately time-consuming because rendering and feature extraction occur inside the fitness-evaluation loop. For ultra-large scenes, memory and compute requirements may increase further. The attack is also explicitly black-box only and does not exploit gradients that could be available in white-box settings. Possible defenses suggested in the paper include methods that do not rely on CNN-feature-based extraction, or methods that robustify watermark extraction against feature homogenization. The authors also suggest surrogate models and advanced black-box optimization strategies such as gradient estimation and natural evolution strategies as possible future improvements to the attack procedure itself (Zeng et al., 10 Aug 2025).
In the security literature around watermarking, GMEA’s primary significance is therefore twofold. First, it demonstrates that current 3DGS copyright protection schemes can be attacked through a universal black-box framework that preserves high visual fidelity. Second, it reframes robustness evaluation for 3D watermarking as a large-scale multi-objective search problem in which detector blindness, not merely visual corruption, becomes a first-class attack objective (Zeng et al., 10 Aug 2025).