How it works

Methodology & measured accuracy

What GPTTrace checks, how it turns evidence into a probability, which models it runs on your device, and how well it actually performs on labelled data — including where it fails.

One model for evidence

Every GPTTrace detector produces a list of evidence items. Each item has a signed strength — positive leans toward AI, negative toward human or authentic — and a weight. The probability shown is a logistic combination: the weighted strengths are added to a baseline and passed through a sigmoid. The weights are fitted by L2-regularised logistic regression on labelled data, regularised toward hand-set starting values so that rare but meaningful signs (an AI disclaimer, a chat-interface citation token) are not erased just because the test set contains few of them. Each data source is weighted equally, so one large dataset can’t dominate.

After fitting, the baseline is shifted so that no more than 5% of human samples reach the “Likely AI” line at 70%. Provenance — a signed C2PA manifest recording AI generation, the IPTC trainedAlgorithmicMedia label, embedded diffusion settings or NovelAI’s alpha-channel record — bypasses the model and settles the verdict, because the file itself declares its origin.

Images

Provenance and metadata are read by GPTTrace’s own parsers: a C2PA reader that walks the JUMBF boxes in JPEG, PNG, WebP, HEIC, AVIF and MP4 files, decodes the CBOR claim and assertions and reads the signer’s X.509 certificate; EXIF, XMP and IPTC via the open-source exifr library; and PNG text chunks including compressed ones. The neural classifier is the Community Forensics ViT-Small model (Park & Owens, CVPR 2025), trained on 2.7 million images from 4,803 generators, run with ONNX Runtime WebAssembly in an int8 version of 22 MB. Images are preprocessed as the model specifies (shortest edge 440, centre crop 384); large images are also scored on four native-resolution crops. Forensic measurements — error-level analysis, the radial spectrum and a lattice-peak detector on a native-resolution crop — run in a Web Worker.

Models we tested and did not ship: a SigLIP-based “AI vs human” image classifier (Apache-2.0, 88 MB) scored an AUC of 0.64 on the same images and lowered the ensemble’s accuracy, so GPTTrace uses the Community Forensics model alone. Two other popular detectors were excluded because their licences forbid commercial use.

image check: AUC 0.938 (cross-validated)

679 labelled samples (399 AI, 280 human), run 2026-10-08. At the “Likely AI” line it caught 64% of AI samples and wrongly flagged 5% of human ones.

SourceTruthSamplesResult at “Likely AI”
gemini-nano-bananaAI4040% caught
midjourney-v6AI4068% caught
midjourney-v5AI4073% caught
flux-devAI4013% caught
flux-schnellAI4048% caught
sdxlAI40100% caught
gpt-imageAI4030% caught
klingAI3997% caught
leonardo-stablecogAI4098% caught
bitmind-imagine-mixAI4080% caught
fullsize-photosHuman4010% wrongly flagged
open-images-photosHuman405% wrongly flagged
lfw-facesHuman400% wrongly flagged
caltech-objectsHuman403% wrongly flagged
coco-photosHuman400% wrongly flagged
ffhq-facesHuman400% wrongly flagged
celeba-facesHuman4018% wrongly flagged

Text

The standard check combines 22 signs derived from Wikipedia’s Signs of AI writing and an era-weighted vocabulary, sentence and paragraph statistics, and a compression-similarity measure using the browser’s deflate implementation. The optional deep scan adds the TMR RoBERTa detector, fine-tuned on the RAID benchmark, run locally after a 120 MB download. Evaluation uses the TextSight 2026 control set (pre-2022 Wikipedia, non-native-English academic writing, public-domain literature and genre-matched output from five 2026 models) and the HC3 corpus. Text numbers are five-fold cross-validated.

standard text check: AUC 0.795 (cross-validated)

2,467 labelled samples (1,258 AI, 1,209 human), run 2026-10-08. At the “Likely AI” line it caught 34% of AI samples and wrongly flagged 5% of human ones.

SourceTruthSamplesResult at “Likely AI”
claude-opus-5AI1250% caught
gpt-4.1AI12100% caught
gpt-oss-120bAI12100% caught
claude-haiku-4-5AI1283% caught
qwen-3.8-27bAI863% caught
pd-literatureHuman3061% wrongly flagged
human-pmc-eslHuman1835% wrongly flagged
human-wikipedia-pre2022Human1558% wrongly flagged
hc3-open_qaHuman714% wrongly flagged
chatgpt-3.5AI120232% caught
hc3-wiki_csaiHuman5586% wrongly flagged

text check with deep scan: AUC 0.91 (cross-validated)

2,467 labelled samples (1,258 AI, 1,209 human), run 2026-10-08. At the “Likely AI” line it caught 56% of AI samples and wrongly flagged 5% of human ones.

SourceTruthSamplesResult at “Likely AI”
claude-opus-5AI1242% caught
gpt-4.1AI12100% caught
gpt-oss-120bAI12100% caught
claude-haiku-4-5AI12100% caught
qwen-3.8-27bAI875% caught
pd-literatureHuman3060% wrongly flagged
human-pmc-eslHuman1831% wrongly flagged
human-wikipedia-pre2022Human1551% wrongly flagged
hc3-open_qaHuman70% wrongly flagged
chatgpt-3.5AI120254% caught
hc3-wiki_csaiHuman5589% wrongly flagged

Audio and video

Audio is decoded at its native sample rate after reading the container header, and measured over the first 60 seconds: spectral band limit relative to the codec’s expected low-pass, exact digital silence, noise floor, pitch micro-variation, loudness spread and steady upper-band tones. Speech-only measurements are skipped for music. Video is checked through its container metadata and Content Credentials, then eight frames are sampled and passed through the image classifier and spectral checks. Audio and video weights are currently hand-set; labelled evaluations for both are being built and their numbers will appear here when they are reliable enough to publish.

Frequently asked questions

Why publish accuracy numbers that aren’t perfect?
Because every detector has error rates, and a tool that hides them invites misuse. Knowing that a check catches most current chatbot text but misses some Claude Opus output, or that music is harder than speech, lets you decide how much a result should count.
What does AUC mean?
The area under the ROC curve: the probability that a randomly chosen AI sample gets a higher score than a randomly chosen human one. 0.5 is a coin flip and 1.0 is perfect separation. It summarises performance across all thresholds.
Why is the threshold set to limit false positives?
Wrongly accusing a person — a student, a photographer, a musician — does more harm than missing an AI file. We set the “Likely AI” line so that at most about 5% of human samples in our test data cross it, and report what that costs in missed detections.
Is the test data public?
The sources are public datasets listed on this page, and the evaluation harness is part of GPTTrace’s code. We don’t redistribute the files themselves because their licences vary.
How often are the numbers updated?
Whenever the detectors, models or weights change, the evaluation is re-run and these tables are regenerated from the results.