Our current evidence
Image Evidence has not published a real-world accuracy benchmark. Our development fixtures exist only to check that the evaluation pipeline runs: their labels do not make them representative of AI-generated images or camera photos, and any metric on them says nothing about performance on real files or current generators. We will not present such numbers as detection accuracy.
What the metrics mean
True-positive rate is the fraction of evaluated AI-labeled samples that receive the likely-AI label. False-positive rate is the fraction of evaluated real-labeled samples incorrectly assigned that label. The uncertainty ratio measures how often the pipeline declines a confident label. Failed files are reported separately. An empty denominator produces no metric rather than a misleading zero. Review the counts and corpus description before interpreting any displayed percentage.
How to build a meaningful corpus
Use independently documented originals with permission and reliable origin labels. Include camera images, screenshots, scans, compressed reposts, generated images from relevant versions, and a separate category for mixed edits. Avoid making all real images come from one device or all generated examples share a resolution. Keep related transformed copies together when separating training and evaluation so duplicates do not make performance appear better than it is.
Measure the trade-off
A narrower uncertain band may raise apparent coverage while increasing harmful false positives. Report threshold choices and uncertainty alongside recall. Break down results by generator and transformation rather than averaging away a weak subgroup. Compare models on the same corpus and preserve the evaluation date. A third-party model card or a vendor's advertised accuracy is not evidence about this application's pipeline, especially when the export or preprocessing differs.
How this page is populated
This public page presents the evaluation method only. Results will be published after a permissioned, representative test set and a real model have been evaluated and independently reviewed. No local fixture metrics are displayed here. Future reports must identify the corpus, model version, date, thresholds, uncertainty, failures, and precise denominators. We will not treat synthetic geometry or mock outputs as evidence of real-world detection accuracy.
What is a false positive in AI image detection?
A false positive is a real, camera-captured image that a detector labels as AI-generated. False positives are especially harmful when a score is used to judge a student's work, a job application or a news photo, so a detector's false-positive rate should be reported on realistic photographs, including compressed and edited ones.
Why not publish one accuracy number?
Accuracy depends on the test set: which generators and versions, how files were transformed and how many real photos were included. One number averages away weak spots and can be inflated by easy examples. Per-generator results, false-positive rates and uncertainty rates with exact counts are more honest, and that is what we will publish.