BoardVision: Motherboard Defect Detection
- The paper introduces BoardVision, a reproducible framework that fuses YOLOv7 and Faster R-CNN using Confidence-Temporal Voting (CTV) for detecting mounting- and wiring-related motherboard defects.
- It leverages the MiracleFactory Motherboard Dataset with 389 high-resolution images and 2,860 annotated defect instances across 11 classes to benchmark performance under realistic conditions.
- The CTV ensemble harmonizes high-speed detection and high-recall error correction, enhancing precision, recall, and stability against imaging perturbations in production QA.
BoardVision is a reproducible framework for robust assembly-level motherboard defect detection that consolidates two leading computer vision detectors—YOLOv7 and Faster R-CNN—through a lightweight, interpretable ensemble strategy termed Confidence-Temporal Voting (CTV). It addresses the underexplored challenge of detecting mounting- and wiring-related defects in assembled motherboards under realistic conditions, providing a practical and deployable end-to-end system that bridges the research-production gap in electronics quality assurance (Hill et al., 16 Oct 2025).
1. Dataset and Task Scope
BoardVision’s empirical foundation is the MiracleFactory Motherboard Dataset, which comprises 389 high-resolution (640×640) RGB images with a total of 2,860 defect instances annotated across 11 assembly-level defect classes. These classes capture a representative spectrum of faults encountered during motherboard production, including missing screws, loose or incorrect fan fittings, surface scratches, and detached connectors.
Classes and Distribution
| Class Name | Instance Count |
|---|---|
| Screws | 806 |
| CPU_FAN_Screws | 685 |
| CPU_FAN_NO_Screws | 326 |
| CPU_fan | 313 |
| No_Screws | 196 |
| CPU_fan_port | 159 |
| CPU_FAN_Screw_loose | 99 |
| Scratch | 95 |
| Incorrect_Screws | 63 |
| CPU_fan_port_detached | 60 |
| Loose_Screws | 58 |
Defect types of interest include missing screws, loose/incorrect mounting, fan mis-wiring, connector detachment, and surface scratches—reflecting real-world assembly line QA requirements.
2. Baseline Detector Performance
BoardVision benchmarks two primary object detectors:
- YOLOv7: a real-time, single-stage CNN detector, trained at 640×640 input resolution, optimized with SGD at a learning rate of 0.01 (batch size ≈ 16).
- Faster R-CNN (ResNet-50+FPN): a two-stage region proposal network with a 0.005 learning rate, identical batch size, and early stopping on validation mAP.
Training and evaluation leverage PyTorch (v2.0), CUDA (11.7), and a GTX 1080 GPU. The validation protocol uses a held-out 45-image test set. In addition, data augmentation at inference simulates robustness factors, including horizontal flips, sharpness modulation (Gaussian unsharp masking), and linear brightness variations.
| Metric | YOLOv7 | Faster R-CNN |
|---|---|---|
| [email protected] | 0.914 | 0.766 |
| [email protected]:0.95 | 0.606 | 0.495 |
| Precision | 0.964 | 0.953 |
| Recall | 0.956 | 0.718 |
| F1-score | 0.960 | 0.819 |
| FPS | 22–25 | 8–10 |
YOLOv7 achieves superior inference speed and precision, though Faster R-CNN identifies defects missed by YOLO, particularly rare or ambiguous cases.
3. Confidence-Temporal Voting (CTV) Ensemble
The core technical contribution is the CTV ensemble, designed to harmonize the high-precision, high-speed characteristics of YOLOv7 with the high-recall, error-correcting nature of Faster R-CNN.
Detection Fusion Methodology
For each frame , let detections be denoted . YOLO and Faster R-CNN detections are greedily matched by class and IoU threshold (). For each matched pair of class :
Fused box coordinates:
Unmatched (solo) detections are accepted based on interpretable rules: (1) high confidence (), (2) superior per-class F1 and thresholded confidence, (3) near-tie in F1 and . Class-wise NMS is lastly applied.
A temporal voting function for sequence-level smoothing is:
Pseudocode Outline
- Load YOLO and FRCNN detections for each frame.
- Match boxes by class and IoU.
- Compute weighted box fusion.
- Apply solo rules for unmatched detections.
- Perform NMS and output.
4. Robustness to Realistic Imaging Perturbations
The BoardVision evaluation protocol stresses resilience under deployment-relevant transformations:
- Flip (horizontal mirror),
- Sharpness Up (unsharp mask),
- Brightness Up/Down (linear adjustment).
Average and standard deviation over these perturbations provide insight into both mean performance and stability.
| Metric | YOLOv7 (mean ± std) | Faster R-CNN | CTV Ensemble |
|---|---|---|---|
| Precision | 0.962 ± 0.006 | 0.954 ± 0.008 | 0.964 ± 0.006 |
| Recall | 0.949 ± 0.009 | 0.713 ± 0.003 | 0.954 ± 0.009 |
| F1-score | 0.958 ± 0.007 | 0.816 ± 0.005 | 0.957 ± 0.006 |
YOLOv7 exhibits minor F1 drops under sharpness increase (to 0.942). Faster R-CNN demonstrates greater sensitivity across all perturbations (mean F1 = 0.816). The CTV ensemble both restores YOLOv7’s recall and reduces F1-score variance by approximately 15%, indicating robust stability under adverse imaging conditions. CTV’s mean robustness score 0 is marginally lower than YOLOv7’s, but its variance is substantially reduced.
5. Operator-centric Deployment: GUI-driven Inspection Tool
BoardVision incorporates an operator-facing PySide6/Qt GUI for real-time inspection. The interface presents three synchronized panes (YOLOv7, Faster R-CNN, and CTV outputs), provides selectable input sources (files, webcam, RTSP streams), and exposes parameters (1, 2, solo thresholds) alongside runtime controls (frame skip, pause/resume). Bounding boxes are color-coded by originating model and decision status; operator corrections are logged for later audit.
The deployment stack features:
- Preprocessing and model loading (CPU/GPU)
- Parallel YOLOv7 & Faster R-CNN inference per frame
- Immediate CTV fusion pipeline execution
- Live update and overlay of fused boxes
- Audit log of operator input
- Modes for both streaming (real-time QA) and file-based (batch) operation
This integration couples model transparency, human oversight, and system usability in a production QA context.
6. Quantitative and Practical Findings
The CTV ensemble achieves the highest overall performance equilibrium: [email protected] = 0.921, [email protected]:0.95 = 0.604, Precision = 0.967, Recall = 0.962, F1 = 0.964, at inference speed comparable to Faster R-CNN alone (8–10 FPS). Notably, rare classes such as No_Screws realize marked F1 improvement (0.926 [YOLOv7] to 0.963 [CTV]). Robustness analysis demonstrates that CTV stabilizes error rates under lighting and sharpness variations, reducing F1 variance by approximately 15%.
Parameter sweeps reveal the system is stable for 3, 4, and 5; more extreme parameterizations trade off between precision and recall.
Findings underscore that ensemble frameworks leveraging explicit, interpretable rules (e.g., high-confidence overrides, per-class F1-weighting) outperform complex black-box meta-learners in this domain, and GUI-driven transparency supports operator trust and auditability.
In sum, BoardVision illustrates a pathway from laboratory detectors and benchmarks to an operational, operator-interpretable QA system for assembly-level motherboard manufacturing, harmonizing speed, precision, recall, robustness, and human-facing deployment (Hill et al., 16 Oct 2025).