Official PyTorch implementation of the paper "Ignoring the Decoy: Exposing and Tackling Forensic Distractions in Image Forgery Localization using Masked Convolutions" (Xander Staelens, Peter Lambert, Glenn Van Wallendael, Hannes Mareen — IDLab, Ghent University – imec).
![]() (a) Forged image with distraction |
![]() (b) Ground-truth mask |
![]() (c) TruFor — no distraction |
![]() (d) TruFor — distraction present (fails) |
![]() (e) Ours — manual mask (recovers) |
![]() (f) Ours — automatic mask (recovers) |
Image forgery localization (IFL) models can be thrown off by forensic distractions: benign visual elements such as logos, captions, and watermarks that occur naturally in real-world images. We show that state-of-the-art IFL models are highly sensitive to these distractions. CAT-Net and TruFor suffer average performance degradations of 15.06% and 55.65%, respectively, when a distraction is present.
To fix this, we introduce masked convolutions: drop-in replacements for standard convolution and pooling layers that simply ignore masked (distracted) regions of the input during inference. They require no retraining meaning existing pretrained weights can be reused directly. Distraction masks can be provided manually, or generated automatically through an iterative, two-step detection process built into the test scripts.
With masked convolutions, performance degradation drops to just 4.15% for CAT-Net and 2.94% for TruFor.
maskedCNN.py # Masked layer implementations (MaskedConv2d, MaskedAdaptiveAvgPool2d/MaxPool2d, MaskedSequential)
CAT-Net/ # CAT-Net adapted to use masked convolutions
TruFor/ # TruFor adapted to use masked convolutions
images/ # Example input images
covers/ # Example distraction masks (PNG, same size as input, white = masked)
create_env.sh # Conda environment setup script
The masked layers are implemented as subclasses of standard PyTorch layers in maskedCNN.py. Each masked layer additionally takes a binary mask as input (shape Bx1xHxW, 1 = masked/ignored, 0 = unmasked), and returns the updated mask alongside its output so it can be passed to the next layer.
To use them, replace the standard layers in a model with their masked counterparts and thread the mask through the forward pass. No retraining, and no changes to the pretrained weights, are required.
- TruFor with masked convolutions — see the
TruFor/folder. - CAT-Net with masked convolutions — see the
CAT-Net/folder.
Modified files:
CAT-Net/lib/models/network_CAT.pyTruFor/TruFor_train_test/lib/models/cmx/builder_np_conf.pyTruFor/TruFor_train_test/lib/models/cmx/encoders/dual_segformer.pyTruFor/TruFor_train_test/lib/models/cmx/decoders/MLPDecoder.pyTruFor/TruFor_train_test/lib/models/cmx/net_utils.py
bash create_env.shThis creates a local conda environment (env/forensic) with PyTorch and all required dependencies. Pretrained weights for TruFor and CAT-Net must be downloaded separately — see the READMEs in TruFor/ and CAT-Net/ for links — and placed in weights/.
Run inference from within TruFor/ or CAT-Net/:
python ./test.py -w [path to weights]Options:
-in,-out— input images (file, folder, or glob) and output folder (default:../images,../output)-cover— manual distraction masks: a single PNG, or a folder with masks matching the input filenames (white = masked region)-idc— enable iterative automatic distraction masking instead of a manual mask-irn— number of refinement rounds (default: 2)-itr— prediction threshold used to build the mask each round (default: 0.5)-idl— mask dilation, in % of image width (default: 5.0)
Example without masking:
python ./test.py -w ../weights/trufor.pth.tar -in ../images -cover None -out ../outputExample with manual masking:
python ./test.py -w ../weights/trufor.pth.tar -in ../images -cover ../covers -out ../outputExample with automatic iterative masking:
python ./test.py -w ../weights/trufor.pth.tar -idc -in ../images -out ../outputIf you use this work, please cite:
@InProceedings{Staelens_2026_WACV,
author = {Staelens, Xander and Lambert, Peter and Van Wallendael, Glenn and Mareen, Hannes},
title = {Ignoring the Decoy: Exposing and Tackling Forensic Distractions in Image Forgery Localization using Masked Convolutions},
booktitle = {Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops},
month = {March},
year = {2026},
pages = {924-932}
}This work was funded in part by IDLab (Ghent University – imec), by Flanders Innovation & Entrepreneurship (VLAIO), by Research Foundation – Flanders (FWO) (1SA9O26N & G0A2523N), and by the European Union.
This work builds on and adapts TruFor (GRIP-UNINA) and CAT-Net (Myung-Joon Kwon). Please refer to their respective licenses in TruFor/ and CAT-Net/ for terms of use of the underlying models and code.





