---
title: Construction-Aware Preprocessing
url: https://www.emergentmind.com/topics/construction-aware-preprocessing
type: topic
---

# Construction-Aware Preprocessing

A plausible unifying definition of **construction-aware preprocessing** is an upstream design strategy in which masking, segmentation, offline randomness generation, input encoding, or auxiliary structural regularization is tailored to the internal construction of a downstream parser, protocol, learner, or reconstruction system rather than treated as a generic front end. Recent work instantiates this idea in log parsing, dealer-assisted secure comparison, compositional generalization on SCAN, memristive reservoir computing, and facade parsing for inverse procedural modeling. Across these settings, the common objective is to move structure-sensitive work earlier in the pipeline so that the downstream stage operates on inputs already aligned with its assumptions [2412.05254] [2602.19604] [2509.20074] [2506.05588] [2604.09260].

## 1. Domain-general pattern

A recurring pattern in the literature is that preprocessing is no longer a neutral formatting step. Instead, it is optimized for the target construction: regex masking is chosen to improve template induction, dealer-generated masks are specialized to polynomial or prefix-AND comparison circuits, pseudo-construction templates are mined to expose productive form-meaning pairings, parity rows are added to match memristor dynamics, and pairwise alignment penalties are introduced to stabilize facade grids. In each case, the upstream operation is justified not by generic cleanliness of the input, but by a specific downstream structural criterion.

| Domain | Construction-aware operation | Reported effect |
|---|---|---|
| Log parsing | Activated, ordered regex masking before statistic-based parsing | Drain FTA increased by 108.9% |
| MPC comparison | Dealer precomputes domain-specific correlated randomness | $1.79\times$ to $19.4\times$ speedups |
| SCAN | Unsupervised pseudo-construction mining and slot-based segmentation | 47.8% on ADD JUMP; 20.7% on AROUND RIGHT |
| Memristive RC | Thresholding, sectioning, 2D projection, parity encoding | Up to $\approx 92\%$ accuracy on MNIST |
| Facade parsing | Alignment loss encouraging grid-consistent boxes | 15–30% reduction in SVD-based regularity metric |

The same pattern also clarifies what construction-aware preprocessing is not. The log-parsing study argues that preprocessing had often been treated as an ad hoc step; the secure-comparison study similarly criticizes dealer-assisted methods that use the dealer merely as a drop-in replacement for conventional preprocessing. Construction-awareness, in this sense, is a rejection of generic offline assistance in favor of preprocessing co-designed with the target algorithmic structure [2412.05254] [2602.19604].

## 2. Log parsing as ordered structural masking

In log parsing, the most explicit formulation appears in the general preprocessing framework proposed for statistic-based parsers such as Drain, IPLoM, LFA, and LogCluster. The framework is a separate first stage composed of four modules: a curated repository of 15 generalizable regular expressions, a regex applicability estimator that scans the first $K$ lines of a new log file with default $K=2000$, an ordered masking engine, and a placeholder normalizer that converts matches to the token `"<*>"`. The end-to-end flow is `raw log → Regex Applicability Estimator → Activated R_i list → Ordered Masking Engine → masked log → statistic-based parser → templates & variables` [2412.05254].

The masking operation is defined tokenwise. For a tokenized message $L=[t_1,t_2,\dots,t_n]$ and activated regex set $R=\{r_1,\dots,r_m\}$,
$$
M(t_i)=
\begin{cases}
\texttt{"<*>"} & \text{if } \exists r\in R \text{ such that } \mathrm{full\mbox{-}match}(r,t_i) \\
t_i & \text{otherwise}
\end{cases}
$$
and the preprocessed message is $P(L)=[M(t_1),M(t_2),\dots,M(t_n)]$. Although implementation uses global regex replacement over raw text, the paper states that this tokenwise view is equivalent. The framework’s central design choice is *ordered masking*: more specific patterns such as MAC, IPv6, and URL are applied before more general ones such as time and path so that overlapping patterns do not mis-mask each other.

Evaluation is reported on Loghub-2k and Loghub 2.0. The framework improves all four parsing metrics for Drain, with average gains across 14 systems of $+0.8\%$ in GA, $+48.4\%$ in FGA, $+35.5\%$ in PA, and $+108.9\%$ in FTA. IPLoM shows $-10.1\%$ GA, $+17.4\%$ FGA, $+67.4\%$ PA, and $+95.0\%$ FTA; LFA shows $-3.2\%$ GA, $+20.8\%$ FGA, $-5.6\%$ PA, and $+59.0\%$ FTA; LogCluster shows $-5.6\%$ GA, $+2.5\%$ FGA, $+277.3\%$ PA, and $+60.0\%$ FTA. Against semantic parsers, Drain without refinement has GA $=0.843$, FGA $=0.554$, PA $=0.468$, and FTA $=0.282$, whereas Drain with the framework reaches GA $=0.847$, FGA $=0.815$, PA $=0.638$, and FTA $=0.581$. Compared to the best semantic parser in that comparison, the paper reports Drain gains of $+28.3\%$ GA, $+38.1\%$ FGA, $-16.1\%$ PA, and $+18.6\%$ FTA. Figures 2–4 further show that the largest FTA and FGA gains accrue on rare and highly variable templates.

This case establishes two durable themes. First, construction-aware preprocessing can improve downstream structure recovery even when some clustering-oriented metric such as GA decreases. Second, selectivity matters: only regexes with at least one match in the sample are activated, which keeps preprocessing fast and dataset-adaptive rather than universally applied.

## 3. Dealer-assisted secure comparison as offline construction matching

In multi-party computation, construction-aware preprocessing is realized through a passive, non-colluding dealer that generates correlated randomness specifically matched to the online comparison construction. The paper formalizes this via an extended Arithmetic Black-Box, $\mathcal{F}_{\mathrm{ABB}^{+}}$, providing ideal functionalities including $\mathcal{F}_{\mathrm{Input}}$, $\mathcal{F}_{\mathrm{LinComb}}$, $\mathcal{F}_{\mathrm{Mult}}$, $\mathcal{F}_{\mathrm{m\mbox{-}Mult}}$, $\mathcal{F}_{\mathrm{Open}}$, $\mathcal{F}_{\mathrm{B2A}}$, $\mathcal{F}_{\mathrm{daBits}}$, and $\mathcal{F}_{\mathrm{edaBits}}$. The dealer acts only in preprocessing and can sample large amounts of correlated randomness in $\mathbb{F}_2$, $\mathbb{Z}_{2^k}$, and $\mathbb{F}_p$, including perfectly consistent decompositions of the same mask in both bitwise and arithmetic form [2602.19604].

For the $\mathbb{F}_p$ LTBits protocol $\Pi_{\mathrm{LTBits}_p}$, the dealer samples random masks $r_i\in\mathbb{F}_p$, distributes their power sequences $\langle r_i^k\rangle_p$ for $k=1,\dots,\ell+1$, and shares polynomial coefficients for
$$
f^{(\ell)}(t)=\prod_{u=1}^{\ell+1}(u-t)=\sum_{k=0}^{\ell+1}\alpha_k t^k,
\qquad
\gamma=((\ell+1)!)^{-1}\in\mathbb{F}_p.
$$
Online, the parties compute
$$
\langle d_i\rangle_p=\langle (a_i-b_i)-r_i\rangle_p,
$$
open all $d_i$ in one round, reconstruct
$$
z_i^k=\sum_{j=0}^k \binom{k}{j} r_i^j (d_i)^{k-j},
\qquad
e_i=\gamma\sum_{k=0}^{\ell+1}\alpha_k z_i^k,
$$
and return
$$
1\{a<b\}=\sum_{i=0}^{\ell-1} e_i.
$$
Only the open step is interactive; the rest is local. The construction therefore achieves constant-round online complexity over $\mathbb{F}_p$.

For $\mathbb{Z}_{2^k}$ MSB extraction, the preprocessing is organized around a prefix-AND or prefix-OR tree with tunable branching factor $n$. The dealer precomputes, for every size-$m\le n$ AND gate, secret shares of all nonempty subsets’ masked ANDs of offline randomness, as well as pre-grouped random vectors for every level of the tree. Each tree level is then evaluated in one round, giving total online round complexity $O(\log_n k)$ and communication dominated by $O(k\log_n k)$ bits.

The paper emphasizes portability: all constructions are black-box over $\mathcal{F}_{\mathrm{ABB}^{+}}$, so they transport to any backend realizing that interface, including SPDZ-style arithmetic sharing, TinyOT, Replicated SS, and Shamir SS. Security is stated for passive adversaries with dishonest-majority $(t<n)$ or honest-majority $(2t<n)$, and the paper states that standard compilers with authentication extend the approach to active adversaries. Empirically, with 3–5 parties in LAN and WAN settings, $\Pi_{\mathrm{MSB}_p}$ achieves up to $19.4\times$ speedup in WAN for small-batch 10-comparison workloads on a 61-bit prime, and $5\times$–$8\times$ on 16–31-bit primes; $\Pi_{\mathrm{MSB}_{2^k}}$ achieves $1.79\times$–$2.78\times$ speedup on 10,000 comparisons by tuning $n$. The main lesson is that preprocessing gains are largest when offline randomness is designed around the online algebraic construction, not merely supplied as generic commodities.

## 4. Pseudo-constructions and compositional generalization on SCAN

For sequence-to-sequence learning on SCAN, construction-aware preprocessing takes the form of unsupervised pseudo-construction mining. The goal is to induce reusable templates with variable slots, such as “`_ around _ twice`”, from unsegmented source sentences. The pipeline has three stages: candidate extraction and masking, misalignment scoring, and beam-search segmentation. Candidate extraction considers all contiguous spans of length up to 4 and all masked variants obtained by replacing one or more non-consecutive tokens with the placeholder “`_`”, with each candidate pattern $P$ assigned a frequency $f(P)$. Misalignment scoring then evaluates whether different fillers for the same pattern preserve target-side structure [2509.20074].

For a candidate $P$, let $\mathcal{S}(P)=\{(s_i,t_i)\}_{i=1}^n$ be the training pairs whose source span matches $P$. For each $s_i$, the nearest neighbor $\mathrm{NN}(s_i)$ is found under Levenshtein distance among the remaining $s_j$, and the target discrepancy is measured by
$$
\Delta_i=\left|\mathrm{len}(t_i)-\mathrm{len}(t_j)\right|.
$$
The misalignment score is
$$
MS(P)=\frac{1}{|\mathcal{S}(P)|}\sum_{i=1}^{|\mathcal{S}(P)|}\Delta_i.
$$
Beam search segments each sentence into top-$K$ candidates or single raw tokens, using the score
$$
\mathrm{score}(Y)=\sum_{k=1}^m \log f(y_k)-\lambda\sum_{k=1}^m MS(y_k).
$$
Surviving patterns with underscores are converted to slot-bearing templates such as `W_1 around W_2 twice`, with explicit mappings from slot tokens back to lexical fillers.

The preprocessing is applied to both source and target. On the source side, a flat sequence such as “jump around right twice and walk twice” becomes a segmented form such as “`W_1 around W_2 twice and W_3 thrice`”, with slot mappings recorded. On the target side, SCAN’s one-to-one primitive lexicon is used to replace corresponding action tokens with the same slot symbols. A standard encoder-decoder Transformer with 4 layers, 4 heads, embedding dimension 256, and feed-forward dimension 1024 is then trained on these slot-augmented sequences, with no architectural changes and no additional supervision.

The reported gains are concentrated on out-of-distribution splits. On ADD JUMP, exact-match accuracy rises from 1.3% for the baseline Transformer to 47.8%; on AROUND RIGHT, it rises from 2.1% to 20.7%. Under subsampling at $k=10$, corresponding to approximately 40% of the original training data, accuracy is 40.7% on ADD JUMP with 5,660 examples (39% of full data) and 9.4% on AROUND RIGHT with 5,423 examples (36% of full data). The paper also reports that removing the misalignment filter drops full-data accuracy by 8–12 points on both splits. This makes the SCAN study a direct demonstration that preprocessing can encode compositional bias without modifying the core seq2seq architecture.

## 5. Hardware-aware encoding in memristive reservoir computing

In memristive reservoir computing, construction-aware preprocessing is defined by compatibility with volatile memristor dynamics. The evaluated preprocessing methods include binary thresholding, sectioning, 2D projection, and parity-based encoding. Binary thresholding maps each grayscale pixel $x_{i,j}\in[0,255]$ to
$$
x_{i,j}^{\mathrm{bin}}=
\begin{cases}
1 & \text{if } x_{i,j}>T \\
0 & \text{otherwise}
\end{cases}
$$
with $T=25$ in the paper. This matches a write–no-write regime in which a ‘1’ produces a voltage pulse and a ‘0’ lets the device decay. The 1D variant feeds each of the $n$ image rows into one memristor; the 2D variant appends the $m$ columns as additional “rows” after a 90° rotation; sectioning splits each row or column into $k$ contiguous segments, increasing the effective number of devices [2506.05588].

The parity-based method adds a second structured signal:
$$
p_{i,j}=x_{i,j}^{\mathrm{bin}}\oplus x_{i+1,j}^{\mathrm{bin}},
\qquad i=1,\dots,n-1.
$$
These parity rows are appended to the original data, yielding $2n-1$ rows for 1D+parity and $n+m+(n-1)$ rows for 2D+parity. The paper’s interpretation is that parity rows highlight contours and edges while remaining sparse, so they add modest write energy relative to the increase in device count. This is explicitly described as construction-aware because it uses only XOR of existing rows, requires only binary pulses, and produces sparse pulses that match the device’s low-energy write regime.

The underlying device model couples preprocessing choices to physical dynamics. The current-voltage relation is
$$
I=(1-w)\cdot \alpha[1-\exp(-\beta V)] + w\cdot \gamma \sinh(\delta V).
$$
Under a write pulse of amplitude $V_p$ and duration $t_p$,
$$
\Delta w=R(w)\cdot t_p\cdot \lambda \sinh(\eta V_p),
\qquad
R(w)=1-\exp(3w)/\exp(3w_{\mathrm{Max}}),
$$
while under no pulse for $t_p$,
$$
\Delta w=(w-w_{\mathrm{Min}})\cdot (1-\exp(-t_p/\tau)).
$$
The approximate write-energy model is
$$
E_{\mathrm{pulse}}\approx C\cdot V_{\mathrm{write}}^2\cdot t_p,
\qquad
E_{\mathrm{image}}\approx N_{\mathrm{devices}}\cdot \langle \#1\mathrm{s\ per\ train}\rangle \cdot E_{\mathrm{pulse}}.
$$

On MNIST, the paper reports $\approx 72\%$ accuracy for 1D without sectioning, $\approx 78\%$ for 1D with $k=4$, $\approx 87\%$ for 2D with $k=6$, $\approx 88\%$ for 1D+parity with $k=4$, and $\approx 92\%$ for 2D+parity with $k=6$. Throughput scales proportionally with $k$, while energy per image scales roughly linearly with device count; parity increases energy sub-linearly because parity streams are sparse. The study also states that too few devices or overly long pulse trains lead to saturation and low accuracy, whereas excessively large $k$ fragments features and can slightly reduce accuracy. Here preprocessing is therefore not only data-dependent but device-dependent.

## 6. Structural regularization in facade parsing

A related upstream strategy appears in facade parsing, where structural assumptions are injected during training rather than through an explicit input-transformation stage. The method augments YOLOv8-small with a lightweight alignment loss that operates on the set of positive bounding boxes selected by YOLOv8’s dynamic task-aligned assigner. For each unordered pair of boxes of the same semantic class, with approximately zero IoU and edge-coordinate differences below a threshold $T$, the method applies an $L_1$ penalty on corresponding edges. On the $x$-axis,
$$
L_{\text{align}^{x,\,ij}}
=
\left|x_1^{(i)}-x_1^{(j)}\right|
+
\left|x_2^{(i)}-x_2^{(j)}\right|,
$$
with an analogous term on the $y$-axis. The aggregate loss is
$$
L_{\text{align}}
=
\frac{1}{\max(n_x,1)}\sum_{(i,j)\in P_x} L_{\text{align}^{x,\,ij}}
+
\frac{1}{\max(n_y,1)}\sum_{(i,j)\in P_y} L_{\text{align}^{y,\,ij}},
$$
and the total objective is
$$
L_{\mathrm{total}}
=
L_{\mathrm{CIOU}}+L_{\mathrm{BCE}}+L_{\mathrm{DFL}}+W\,L_{\mathrm{align}}.
$$
At inference, nothing changes: the system retains the standard confidence threshold 0.25 and standard NMS from YOLOv8 [2604.09260].

The paper’s geometric prior is that elements of the same semantic class on the same architectural line should share identical left/right or top/bottom edges, subject to a hard horizon $T$ beyond which alignment is not enforced. This horizon is intended to prevent over-alignment of boxes that belong to genuinely different rows, columns, or scales. The use of thresholded pair selection makes the regularizer lightweight and local, but the cumulative effect over many pairs is a global facade grid in which windows and related elements line up in $x$ or $y$.

Experiments are conducted on the CMP facade dataset of Tyleček and Šára (2013), consisting of 606 original street-view images and 689 manually cropped facades, with an 80%/10%/10% random split. The model is YOLOv8-small, anchor-free, pretrained on COCO, fine-tuned for 200 epochs in the Ultralytics PyTorch codebase, with checkpoint selection by highest validation mAP@0.5. The regularization sweep covers $W\in\{0.0,0.1,0.3,0.5,1.0\}$ and $T\in\{6,7,9,10,12\}$. The best trade-off is reported at $T\in\{9,10,12\}$ and moderate $W\approx 0.3$–$0.5$. At $W=0.5$, $T=9$, the SVD-based regularity metric drops to approximately 80–85% of baseline, corresponding to a 15–20% gain in regularity, while mAP decreases by only approximately 0.02–0.04. Pushing $W\to 1.0$ lowers the SVD metric further to approximately 70% but also drops mAP to approximately 0.57. Small $T$ values of 6 or 7 can over-constrain boxes, and for large $W$ the SVD metric can increase beyond 100%, meaning regularity is harmed.

The paper reports qualitative recovery of windows behind trees or in image margins, improved horizontal alignment under residual perspective skew, and better size consistency for windows partially occluded by balconies. It also reports failure cases: missing very small or very large windows and over-stretching or shrinking when real facades contain mixed window types in one row. For downstream inverse procedural modeling, the significance is that early structural regularization makes the bounding-box layout a more reliable scaffold for grammar induction and procedural fitting, especially for systems such as FaçAID that assume a near-perfect grid of elements.

## 7. Trade-offs, misconceptions, and broader implications

Several misconceptions recur across these literatures. One is that preprocessing is merely a convenience layer. The log-parsing study explicitly argues that lack of understanding of preprocessing can hinder optimal parser use, while the secure-comparison study argues that dealer assistance should not be treated as a drop-in replacement for conventional preprocessing. Another misconception is that better upstream structure must uniformly improve all downstream metrics. The evidence is mixed: in log parsing, aggressive masking may slightly harm GA while improving FGA and FTA; in facade parsing, stronger alignment improves regularity but eventually causes a steep mAP drop; in memristive RC, more sections or more devices help only up to the point where features become fragmented; and in SCAN, templates with highly variable target length receive higher misalignment scores and are down-weighted [2412.05254] [2602.19604] [2604.09260] [2506.05588] [2509.20074].

A second shared theme is selective rather than universal preprocessing. The log framework activates only regexes observed in the first 2,000 lines. The SCAN method keeps only high-quality templates that survive beam-search segmentation and frequency thresholds. The facade method aligns only box pairs of the same class whose edges are already close under threshold $T$. The MPC constructions precompute exactly the randomness needed for a polynomial comparator or an $n$-ary AND tree. The memristive RC study advocates starting from minimal 1D binary thresholding and adding sectioning, 2D projection, or parity only as required by the device budget and accuracy target.

A plausible implication is that construction-aware preprocessing is best viewed as **assumption management**. It externalizes structural assumptions that would otherwise need to be learned implicitly or paid for during online computation. In log parsing, the assumptions concern variable-token regularities; in MPC, algebraic structure and interaction patterns; in SCAN, reusable slot-bearing templates; in reservoir computing, device-compatible temporal statistics; and in facade parsing, near-grid architectural regularity. The empirical record in these papers suggests that when those assumptions are correct, upstream tailoring can substitute for heavier architectural intervention, reduce online interaction, or simplify downstream reconstruction. When they are incorrect or over-applied, the same mechanisms produce over-masking, over-alignment, unnecessary offline cost, or degraded generalization.

Source: https://www.emergentmind.com/topics/construction-aware-preprocessing