Kennen-i: Input Crafting in Machine Learning
- Kennen-i is a spectrum of techniques that craft minimally perturbed inputs to alter model predictions while preserving input validity.
- It employs methods such as gradient-based attacks, frequency decomposition, and structured knowledge injection to achieve controlled outcomes.
- Empirical results across vision, text, and sequence tasks highlight its high misclassification rates and improved performance in multi-hop QA tasks.
Kennen‐i (Input‐Crafting) encompasses a spectrum of algorithmic techniques for constructing carefully modified input instances—perturbed at minimal cost—to elicit controlled or adversarial outputs from machine learning models. The term spans adversarial attacks, in-context editing, and knowledge injection, across modalities including vision, text, and sequence modeling. The paradigm is rooted in the analysis of model decision boundaries and feature representations, exploiting local or global sensitivities to systematically alter input and thereby induce model error or controlled behavior. Techniques used in Kennen‐i include both direct optimization in the input space and structurally informed manipulations (e.g., word-level constraints, frequency domain decompositions), with empirical efficacy evaluated across diverse tasks such as multi-hop question answering, image recognition, and sequence classification.
1. Formalizations and Core Objectives
The input-crafting problem is canonically defined as follows: for a fixed classifier or predictor , an original input , and a ground-truth or desired output (or for targeted attacks), the goal is to produce a perturbed input such that (untargeted) or (targeted), while minimizing some cost metric and preserving input validity. This generalizes across domains—continuous (vision), discrete (text), and sequential (RNNs)—necessitating domain-adaptive strategies (Samanta et al., 2017, Papernot et al., 2016, Anshumaan et al., 2020, Wang et al., 31 May 2025).
Key constraints include:
- Semantic similarity: In natural language, adversarial examples must remain grammatical and close in meaning to .
- Imperceptibility: In images, or 0 must be below a visual threshold.
- Task specificity: In in-context knowledge editing, only the targeted knowledge should be updated, preserving the consistency of reasoning paths (Wang et al., 31 May 2025).
Optimization formulations typically blend norm constraints with model-driven loss objectives, as in: 1 for task loss 2, often further regularized or subjected to input domain projection.
2. Algorithmic Strategies and Methodological Variants
2.1 Gradient and Saliency-Based Perturbation
In continuous domains (e.g., images, RNN-encoded sequences), the Fast Gradient Sign Method (FGSM) efficiently computes perturbation directions: 3 For RNNs, gradients are computed through backpropagation through time at each timestep (Papernot et al., 2016).
Text domain attacks adapt saliency concepts: token importance is quantified via “leave-one-out” contribution or gradient surrogates,
4
Tokens are greedily edited—deleted, swapped with synonyms or genre-contrasted terms, or minimally inserted—subject to part-of-speech consistency, leveraging class- and subcategory-keyword statistics for maximal impact (Samanta et al., 2017).
2.2 Universal and Data-Free Objectives
Generalizable data-free adversarial objectives (GD-UAP) seek a perturbation 5 maximizing neuron activations network-wide: 6 where 7 are the activations at 8 layers. Variants incorporate priors, e.g., sampling pseudo-data from a Gaussian with known input statistics: 9 This approach yields image-agnostic, black-box transferable, and highly generalizable perturbations, with fooling rates on ImageNet classifiers in the 49–92% range depending on data regime (Mopuri et al., 2018).
2.3 Frequency Domain and Structural Decomposition
WaveTransform adversarial attacks decompose input images via discrete wavelet transforms (DWT), generating perturbations constrained to specific frequency subbands (LL, LH, HL, HH) and reconstructing perturbed images meeting 0 bounds:
- Targeted gradient updates are performed in wavelet-coefficient space only for chosen bands, guided by cross-entropy loss.
- The attack can focus on low-frequency (coarse) or high-frequency (edge/detail) subbands, with all-band perturbation yielding near 100% misclassification in white-box settings (Anshumaan et al., 2020).
2.4 Masked Reasoning and Knowledge Path Decoupling
For LLMs, DecKER introduces a two-phase, input-crafting approach that separates reasoning path planning (generation of masked, type-annotated chains-of-thought) from knowledge injection (fact filling and conflict validation):
- The masked reasoning path is generated via prompting with “[MASK]” tokens and type hints.
- Fills for masks are determined by retrieval from edited fact memory and/or LLM validation, using score thresholds 1 for rapid conflict detection.
- Two-round evaluation with predictive entropy and type-matching metrics stabilizes performance without parameter updates, correcting the >80% drop in multi-hop QA accuracy observed in earlier, entangled ICE approaches (Wang et al., 31 May 2025).
3. Domain-Specific Input-Crafting Techniques
| Domain | Core Crafting Principle | Example Techniques |
|---|---|---|
| Vision (images) | Direct pixel and frequency subband perturbation; universal noise | FGSM, PGD, GD-UAP, WaveTransform |
| Text | Word saliency ranking, POS-preserving edits, genre-keyword leveraging | Greedy token-level swaps/insertions, contribution scoring |
| Sequential | Temporal and feature-wise gradient attack, targeted output saliency | RNN-adapted FGSM, JSMA, continuous-time step perturbations |
| LLM/Reasoning | Chain-of-thought masking, independent knowledge injection | DecKER masked paths, retrieval-validation-filling |
Each setting imposes unique constraints: images require imperceptibility, text and sequence models demand linguistic or semantic plausibility, knowledge editors must uphold reasoning path consistency.
4. Evaluation Protocols and Empirical Findings
- Vision: Perturbations crafted by GD-UAP and WaveTransform show transfer and fooling rates (2–3) across architectures and tasks (classification, segmentation, depth estimation). WaveTransform achieves 4 white-box error when all subbands are perturbed; LL (low-frequency) perturbations exhibit stronger black-box transfer than only high-frequency perturbations (Anshumaan et al., 2020, Mopuri et al., 2018).
- Text: Saliency-based greedy token editing with genre-keyword enlargement reduces CNN-based sentiment and gender classifier accuracy to below 60%, with high semantic similarity maintained (5 spaCy metric 6) (Samanta et al., 2017).
- Sequential: Papernot et al. demonstrate that on LSTM sentiment models, 7 of token edits flip all training instances; sequence-to-sequence RNNs require alterations to a small subset of input steps to induce large output deviations (Papernot et al., 2016).
- Knowledge Editing: DecKER improves multi-hop QA task accuracy by 8 percentage points over standard ICE, achieving reasoning-framework similarity of 9 (pre-/post-edit), and operates efficiently under hybrid retrieval + LLM validation schemes (Wang et al., 31 May 2025).
5. Algorithmic Pseudocode and Workflow Patterns
Representative pseudocode for input-crafting algorithms aligns with the following high-level steps:
- Saliency/Contribution Calculation: Quantify influence of input components (tokens, pixels, features).
- Perturbation Selection: Rank order by score; select targets for editing or noise injection.
- Candidate Pool Construction: For text, build substitute sets (synonyms, genre-opposite keywords); for vision, decompose into subbands; for knowledge, retrieve and validate candidate facts.
- Edit Application: Apply edits under domain constraints (POS, 0-norm, type-matching).
- Validation and Selection: Evaluate crafted samples via model confidence reduction/entropy, semantic similarity, or type-verification.
The DecKER algorithm instantiates this pipeline for knowledge editing as follows (Wang et al., 31 May 2025):
- Generate 1 masked reasoning paths;
- For each, iteratively fill [MASK]s via hybrid retrieval/validation;
- Two-stage selection: filter high-entropy plans, then select max type-match candidate.
6. Limitations, Defenses, and Open Directions
- Greedy input-crafting (especially in text) is heuristic and may not achieve global optimum; very short texts and non-token cues (e.g., punctuation) are problematic (Samanta et al., 2017).
- RNN and LLM defenses (adversarial training, robust optimization, input denoising) are extensions of feed-forward paradigms, but with many open practical questions regarding efficacy (Papernot et al., 2016).
- Frequency decomposition attacks remain potent even under kernel smoothing and adversarially trained defenses—a key challenge for robustification in vision models (Anshumaan et al., 2020).
- Decoupling reasoning and fact-injection in LLMs is crucial for robust in-context knowledge editing, highlighting a broader need for compositional control interfaces in large models (Wang et al., 31 May 2025).
A plausible implication is that future research in Kennen-i will increasingly integrate domain structure, multi-modal constraints, and hybridized retrieval/model-in-the-loop mechanisms to enhance both the efficacy and control of input-crafting algorithms, particularly as models become more capable and context-dependent.