← ALL CASE FILESTHE INTERNET SAID WHAT?
CASE002
CASE 002

AI Detection

INTERNET CLAIMS. ACTUAL EXPERIMENTS.

Can GPTZero Really Detect AI Writing With 99% Accuracy?

GPTZero says it can distinguish AI-generated text from human writing with 99% accuracy. CASE 002 will test the headline claim on clean human/AI samples, then stress it with edited, mixed, and style-imitated text.

Published
  • ai
  • ai-detection
  • writing
  • education
  • generative-ai
INITIAL READ 3.2/10
CASE 002 / AI Detection

PRELIMINARY ASSESSMENT

INITIAL READ / BEFORE TESTING

Preliminary score. An initial reading of the claim, before the test.

PRELIMINARY BS METER
3.2/10

THE INTERNET SAID WHAT? — CASE 002

THE 99% DETECTOR

The Claim

GPTZero currently says its AI detector has a 99% accuracy rate when distinguishing AI-generated text from human writing.

That is a strong, specific claim. It is also unusually testable.

GPTZero’s own current benchmarking reports performance close to or above that level on several clean benchmark sets. Independent testing has also found strong performance on ordinary AI-generated text, although harder conditions such as style imitation and edited AI text can reduce detection rates.

The question for CASE 002 is simple:

If we give GPTZero writing where we already know who—or what—wrote it, does it really get the answer right 99% of the time?

Preliminary BS Meter

3.2 / 10

This starts much lower than CASE 001.

There is credible evidence that GPTZero is genuinely good at this task. The suspicious part is not that AI detection is obviously fake; it is that “99% accurate” is easy to hear as a universal guarantee, when actual performance can depend on the model, writing domain, document length, editing, mixed authorship, and deliberate attempts to make AI text look human.

The working hypothesis is:

The detector is probably very good on clean human-vs-AI text, while the 99% headline becomes less universal once the writing gets messy.

Why We’re Suspicious

1. “99% accuracy” compresses several different questions into one number

Overall accuracy can hide the errors that matter most.

For example:

  • How often does human writing get falsely accused of being AI?
  • How often does AI writing escape detection?
  • Does accuracy change across fiction, technical writing, essays, blogs, and casual writing?
  • Does performance change with different AI models?
  • What happens when a human edits AI text?
  • What happens when AI imitates a specific human style?

A detector can have excellent overall accuracy while still having edge cases worth caring about.

2. False positives carry asymmetric consequences

If an AI-written paragraph is missed, the detector made an error.

If a genuinely human essay is flagged as AI and someone is accused of cheating, the same statistical error can carry a much larger real-world cost.

CASE 002 will therefore report human false positives separately, not bury them inside one overall accuracy number.

3. Independent testing shows clean detection can be excellent while harder cases degrade

Recent independent benchmark data has shown GPTZero performing extremely well on direct AI generations and verified pre-ChatGPT human writing, while missing more AI text when models are explicitly prompted to imitate a human author’s style.

That does not make the detector bad. It makes the conditions behind the headline important.

4. The detector keeps changing

GPTZero updates its models over time.

That means this case must record:

  • the test date,
  • the detector/model version when available,
  • the exact sample set,
  • the classification rules used.

The result will describe the version we actually test—not every past or future version.

What Would Change Our Minds

The preliminary BS Meter should move down if GPTZero:

  • scores at or above roughly 99% on the clean human-vs-AI benchmark,
  • keeps the human false-positive rate at or below about 1%,
  • performs consistently across multiple writing domains and current AI models,
  • remains strong under the separate stress-test conditions.

The BS Meter should move up if:

  • clean human-vs-AI accuracy falls materially below the advertised level,
  • human false positives exceed the advertised low error rate,
  • performance varies dramatically by domain or model,
  • modest editing/style imitation causes a large collapse,
  • the headline result depends heavily on favorable sample selection.

We will report the clean benchmark and stress tests separately so harder adversarial samples do not unfairly redefine the company’s simpler headline claim.

The Test

Phase A — Test the headline claim

Build a 200-document ground-truth dataset:

  • 100 verified human-written samples
  • 100 direct AI-generated samples

Use five writing domains:

  1. student/academic-style essays
  2. technical/scientific writing
  3. blogs/opinion writing
  4. fiction/creative writing
  5. casual/informal writing

Target approximately 250–500 words per sample so the detector has enough text to evaluate.

Human samples

Human samples must have strong provenance.

Preferred sources:

  • writing created before the public release of ChatGPT,
  • the research team’s own known pre-AI writing,
  • appropriately licensed/public-domain corpora,
  • documented human benchmark datasets.

Do not use text with uncertain authorship.

AI samples

Generate AI samples using multiple current frontier models.

Balance:

  • model,
  • writing domain,
  • prompt style,
  • approximate document length.

Save:

  • model name/version where available,
  • prompt,
  • raw generated text,
  • generation date.

Primary measurements

For the 200 clean samples calculate:

  • overall accuracy
  • AI recall / true-positive rate
  • human false-positive rate
  • AI false-negative rate
  • confusion matrix

The primary headline comparison is GPTZero’s 99% human-vs-AI accuracy claim.

Pre-registered interpretation

For this Mythadis test:

  • ≥99% clean accuracy: headline survives our clean benchmark
  • 97–98.9%: detector is very strong, but the blanket 99% headline did not reproduce
  • <97%: headline is materially overstated on our sample

False positives will also be called out independently, regardless of the overall score.

These are editorial interpretation bands for this experiment, not universal scientific thresholds.


Phase B — Stress Test

Run a separate 100-sample stress set.

Suggested mix:

  • 30 AI texts prompted to imitate a known human writing style
  • 30 AI texts lightly edited by a human
  • 20 mixed human + AI documents
  • 20 human texts polished/revised with AI assistance

Report this separately from Phase A.

The purpose is not to move the goalposts on the 99% clean classification claim. It is to answer the more useful real-world question:

How quickly does the detector’s certainty fall apart when authorship stops being clean?

Metrics:

  • classification accuracy by stress category
  • AI miss rate
  • human false-positive rate
  • mixed/assisted classification behavior
  • difference versus the clean benchmark

Reproducibility Rules

Before testing:

  1. Freeze the sample dataset.
  2. Hash each source document.
  3. Record provenance for every sample.
  4. Record detector date/version if exposed.
  5. Define how GPTZero’s Human / Mixed / AI-style outputs map to our scoring.
  6. Do not edit samples after seeing detector results.
  7. Store raw detector output where permitted.
  8. Publish aggregate results and representative examples.

If API access is practical, use it for repeatability.

If the dashboard must be used manually, record every result and preserve screenshots/logs.

What Happened

TEST NOT YET RUN

CASE 002 is currently in the experiment-design stage.

No final accuracy result has been assigned.

The Verdict

PENDING

No final BS Meter will be assigned until the benchmark has been run.

The Receipts

Primary claim and technical context:

Independent/reproducible context:

The independent benchmark’s pinned May 2026 GPTZero model produced:

  • 0 false positives among 495 verified human passages,
  • about a 1% miss rate on direct AI generations,
  • about an 11% miss rate on style-imitation AI generations.

That result is part of why CASE 002 starts with a relatively low BS Meter rather than assuming the detector is nonsense.


VIDEO: COMING

FINAL BS METER: PENDING

VERDICT: TEST IN PROGRESS

About the BS Meter: The BS Meter is a Mythadis editorial opinion based on the evidence available at publication. It is not a legal or regulatory finding and does not, by itself, allege fraud, criminal conduct or intentional deception. Editorial Standards

BROADCAST PLACEHOLDER

THE INVESTIGATION.
ON SCREEN.

External video platforms carry the broadcast. This case file keeps the investigation and its receipts together.

No video is attached to this investigation.