Linda-Pro: local English AI-text detector (stylometry + 2 DeBERTa) - looking for hard test texts

Ensemble of stylometry + two fine-tuned DeBERTa models, scored in ~300-word windows, with a per-sentence heat map for mixed authorship.

Results on the Chicago Booth benchmark (1% false positives): plain AI 99.7%, after a StealthGPT humanizer 85%. On generators the models never saw (889 texts): 92.6% flagged. Weak spots published too: TOEFL exam essays 8.8% false positives (1.1% in precise mode), humanizers, short texts.

Weights (gated, free for research): Lindarixon/Linda-Pro · Hugging Face
Code and evaluate.py for your own texts: GitHub - Lendarixon/Linda · GitHub

I’m looking for adversarial and non-native-writer texts to break it. DM me if you want a hosted demo with file upload.

Publishing the weak spots up front is the part that earns trust — rare, and
it makes the hard-test ask real.

One question on the negative class: the 889 unseen-generator texts stress-test
the machine side, but the 8.8% TOEFL false-positive number points at the other
side — human essays getting flagged. Do the negative sets include
non-native-writer human texts at scale, or is that where most of the next
break round should aim?

If the operating point stays at 1% FP, the contrast class that decides whether
it’s usable is humans who write like the weak spot — not the humanizers. I’d
push there: take the flagged TOEFL essays, measure where the stylometry
features separate them from the humanizer outputs, and publish that boundary.
That’s the one number a skeptic will ask for, and right now it’s the hole in
an otherwise strong bench.

One line of proof: a machine can now put a receipt behind a claim like this.
A real one, free to verify — request + payment + signed outcome, all on the page:

That’s a $0.25 compute-spot call settled on Base, signed eip191, free to re-check
forever. Authority you can rotate without trusting the author.

Thanks, appreciate the thoughtful critique!

To answer the negative class question: non-native writing is actually already in our evaluation and calibration at scale, not just the Liang et al. TOEFL set:

  • ELLIPSE (1,500 non-native student essays): 0.7% FP (0.1% in precise mode)
  • PELIC (1,000 adult ESL essays): 0.8% FP (0.3% in precise mode)
  • W&I (280 adult ESL essays): 1.1% FP (0.0% in precise mode)

The 8.8% outlier is specifically on the classic Liang et al. TOEFL benchmark (91 short, formulaic 5-paragraph essays). That specific failure mode is precisely why we added the precise operating mode: DeBERTa flags the rigid formulaic phrasing as AI, but requiring stylistic agreement from the CPU stylometry model drops TOEFL false positives down to 1.1% without cratering detection on standard AI.

Your point about mapping the feature boundary between false-positive TOEFL essays and humanizer outputs is spot on. I’ll extract the stylometry feature vectors (burstiness, function word n-grams, perplexity variance) across those flagged TOEFLs vs. StealthGPT/HumanizerBench and publish the separation plot. That’s indeed the cleanest way to prove it’s not just collapsing ESL into “humanized slop”.

(And neat concept with the Agent Embassy outcome receipts for verifiable machine execution!)

Your numbers seat it: three corpora under 1.2%, TOEFL at 8.8%. The outlier is the question, not the footnote.

True split: run each humanizer per model — several hard reads, one convergence. Post the split rate on humanizer text — derive the honest bound from it.

Three sets, not two: flagged TOEFL, humanizer runs, clean ESL. If the humanizer ink lands on the flagged TOEFLs, that collapse is data too.

Next: the plot you pledged, same essays through both machines.