Start with representative images
Select images that reflect the dimensions, textures, colors and content your platform delivers. Include difficult examples rather than relying only on a small set of visually convenient samples.
Reproduce the distribution path
Apply the transformations your users and downstream platforms actually introduce: format conversion, compression, resizing, cropping or screenshots. Keep original files and transformation settings so results can be reproduced.
Measure visual quality and identifier recovery
Inspect the watermarked image and compare it with the original. Record whether the detector recovers the exact expected ID after each transformation. Report failed and inconclusive attempts, not only successful ones.
Include negative controls
Test images that have not been watermarked, and images from outside the test set. Evaluate false detections and unexpected identifiers. Avoid interpreting a detector confidence value as a direct probability of a leak or an origin claim.
Define an operational decision
Decide what your team will do with a recovered ID, how it will review supporting records and when it will escalate to human review. Keep results tied to the tested configuration, and repeat the evaluation when your pipeline changes.