By May 2026, three frontier AI models dominate professional and academic writing: OpenAI's GPT-5.5, Anthropic's Claude 4.6, and Google's Gemini 3.1. Each ships with marketing claims about how human its outputs feel. Each is used millions of times a day for the kind of writing that ends up in essays, reports, and content libraries. And each leaves a different statistical fingerprint that detection tools can β or sometimes can't β catch.
We ran 600 fresh samples across the three models and pushed each through the same multi-model ensemble detector. The results upend a few common assumptions about which model is hardest to detect β and clarify why the answer matters more in 2026 than in any previous year.
How We Tested
We generated 200 outputs per model β 50 each across academic essays, marketing copy, technical explainers, and creative fiction. Each generation used a default system prompt with no instructions to evade detection. We then created adversarial variants for each: a humanizer-processed version and a lightly edited version (15–20% words manually rewritten). Total: 600 baseline samples, 1,200 adversarial samples.
Every sample was evaluated by Plagly.ai's Agentic Council ensemble (perplexity, burstiness, stylometry, fingerprint, discourse) and cross-checked against the latest classifier from a major commercial single-model detector. Results below report Plagly's detection rate.
Headline Results: Raw Output
On default outputs without adversarial editing, the three models cluster surprisingly close together β but the small differences are revealing.
Detection accuracy on raw outputs
GPT-5.5 (raw): 94% caught. Claude 4.6 (raw): 88% caught. Gemini 3.1 (raw): 92% caught. On default outputs with no adversarial prompting, all three are reliably detected by ensemble tools β but the gap between Claude 4.6 and the others is real.
Why Claude 4.6 leads on stealth
Claude 4.6 was trained with a strong focus on register variation and human-like prose rhythm. The model varies sentence structure more aggressively than its peers, uses fewer formulaic transitions, and shows more genuine commitment to a position in argumentative writing. The result is text that scores closer to human prose on the perplexity and burstiness axes that single-model detectors emphasize.
Why GPT-5.5 is the easiest to spot
Despite being the most heavily optimized for fluency, GPT-5.5 retains the strongest top-down structural patterns of the three. Its training emphasized reliability and balance, which produces a recognizable argument shape: introduce, present multiple perspectives, conclude with a measured synthesis. Detectors that model paragraph- and document-level discourse pick this up consistently, even when sentence-level surface features pass.
Where the Models Diverge: Per-Feature Breakdown
The headline accuracy numbers hide bigger differences in which features trigger detection. Looking at per-classifier scores tells a more useful story.
Discourse signals
GPT-5.5 has the strongest discourse fingerprint of the three. Its preference for balanced arguments, its over-use of certain transitional phrases, and its tendency to pre-summarize make it the easiest of the three frontier models to spot at the paragraph level β even when individual sentences pass single-model detectors.
Stylometric signals
Gemini 3.1 leads on stylometric tells. Its outputs show distinctive patterns in adverb frequency, passive voice construction, and clause-level rhythm. These features are deeply baked into how Google's training data was curated and have proved difficult for the model to suppress without losing fluency.
Perplexity signals
Claude 4.6 has the cleanest perplexity profile β its outputs sit closest to human-written baselines. This is what makes it hardest for single-model perplexity detectors to catch. But Claude 4.6 still has measurable fingerprint and discourse features, which is why ensemble detectors maintain accuracy on it where simpler tools fail.
Test all three models against Plagly's ensemble
Paste GPT-5.5, Claude 4.6, or Gemini 3.1 output into Plagly.ai's free AI detector. See the per-model breakdown β perplexity, burstiness, stylometry, fingerprint, discourse β and which signal triggered the score.
Try Plagly FreeAdversarial Conditions: When Things Get Harder
The bigger gap between models shows up under adversarial conditions. After running each model's output through a humanizer, we re-tested:
- GPT-5.5 + humanizer: 72% caught.
- Claude 4.6 + humanizer: 64% caught.
- Gemini 3.1 + humanizer: 70% caught.
Claude 4.6's lead under raw conditions widens further once a humanizer enters the picture β but even at the most adversarial setting, the majority of Claude 4.6 output is still flagged. The notion that any of these models can reliably evade modern ensemble detectors does not hold up to the data.
Why This Matters in 2026
The model wars accelerated through late 2025 and 2026, with OpenAI shipping GPT-5.5 in a bid to leapfrog Claude's stealth advantage, Anthropic responding with Claude 4.6's prose-rhythm improvements, and Google pushing Gemini 3.1 to close the fluency gap. From a detection standpoint, the practical implications are converging.
Single-model detectors are increasingly unreliable
Across all three models, single-classifier detectors are the most exposed. A perplexity-only tool that worked well on GPT-3.5 now miscategorizes Claude 4.6 output 50–60% of the time. A burstiness-only tool fails similarly. Multi-model ensembles maintain accuracy because no single model needs to do all the work.
Continuous retraining matters more than any single algorithm
Plagly.ai's Agentic Council retrains its fingerprint module continuously on fresh outputs from each major frontier model. Detection accuracy on Claude 4.6 has climbed roughly six percentage points since launch as the model's stylistic patterns became better represented in training data β a pattern that's repeated with every new model release.
What's Next: Claude 5, GPT-6, Gemini 4
All three labs have publicly signaled next-generation releases through 2026 and into 2027. Based on the GPT-5.5/Claude 4.6/Gemini 3.1 cycle, here's what to expect: another temporary detection accuracy dip across all tools at each major release, faster recovery for ensemble products, and a continued widening gap between modern multi-model detectors and legacy single-classifier products.
The longer-term arms race favors detection β slightly. Each generation of LLM reduces shallow stylistic tells but cannot easily eliminate the deeper structural patterns that come from being trained as a next-token predictor on internet-scale data. Those patterns are what well-designed detectors learn to read, and that's the layer that survives model updates.
Run a real-world model showdown
Have outputs from GPT-5.5, Claude 4.6, or Gemini 3.1? Plagly's free detector breaks down which signals each model triggers β and shows you why the same content reads differently to ensemble vs single-model tools.
Run a Free Detection TestThe Bottom Line
Claude 4.6 is currently the hardest to detect of the three frontier models β but only by a few percentage points, and only when measured against single-model detectors. Multi-model ensembles still flag the majority of all three models' outputs, and the gap closes further with each round of detector retraining. The right question for 2026 is not “which model can I use to evade detection?” but “which detector can I trust to give me an accurate read?” The answer to that question β ensemble, multi-model, continuously retrained β has never been clearer.
