The short answer
AI image detectors can help sort large numbers of images, but none is accurate enough to be the only basis for a decision. Performance depends on the generators, file types and edits in a detector's test data, and real-world images are messier than most benchmarks. Treat a score as one piece of evidence, and look for provenance records and context before drawing conclusions.
Why detectors miss new generators
Detectors learn statistical traces of the generators they were trained on. Research by Ojha, Li and Lee showed that detectors trained on GAN images often failed to recognize images from newer diffusion and autoregressive models, classifying many of them as real. Each new model release can open a blind spot until a detector is retrained and re-evaluated.
False positives and the base-rate problem
A false positive is a real photo labeled AI-generated. Even a small false-positive rate matters when most images are real. Suppose 1% of the images in a collection are AI-generated, and a detector catches 99% of them while wrongly flagging 1% of real photos. Among 10,000 images it would flag about 99 AI images and about 99 real ones, so roughly half of its flags would be wrong. Ask for false-positive rates measured on photos like yours.
What compression and edits do
Screenshots, resizing, cropping, filters and social-media compression change the fine details that detectors rely on. A score on a re-shared copy can differ from a score on the original, in either direction. Edited photos, where only part of an image is synthetic, are especially hard for a single overall score to describe.
How to read an accuracy claim
A meaningful claim names the test set, the generators and versions covered, how files were transformed, the number of real and AI images, and the false-positive rate as well as the detection rate. Be cautious of a single headline percentage, results only on easy or unedited images, and tests run only by the vendor. Image Evidence publishes its evaluation method and no accuracy number until a representative test is complete.
Using a detector responsibly
Combine any score with the original file, provenance checks such as Content Credentials and SynthID, and the claim attached to the image. Let a person review uncertain cases, record what remains unknown, and give creators a chance to explain their process. Never use a detector score alone to punish a student, reject an applicant or accuse someone publicly.
Which AI image detector is the most accurate?
There is no stable answer, because accuracy depends on the images you test and changes as new generators appear. Prefer tools that publish independent or reproducible evaluations with per-generator results and false-positive rates, and test them on examples similar to your own before relying on them.
Can AI image detectors be fooled?
Yes. Ordinary transformations such as compression or resizing can change scores, and newer generators may not resemble a detector's training data. This is why a detector result should never be treated as proof on its own. Image Evidence does not provide or explain detector evasion techniques.
Should teachers use AI image detectors?
Only as a prompt for conversation, not as evidence of misconduct on its own. False positives are a real risk, and students may have legitimate reasons for editing images or using generative tools. Ask about the process, request drafts or source files, and follow your institution's policy.