I fed several detectors the same set of AI-involved samples, including human drafts that had later been edited with AI. Then I compared how often each service flagged them. That method matters, since it uses a broader definition of AI involvement than some people will accept.
What the numbers showed
Clever AI Detector flagged 96.7% of the AI-involved texts, narrowly ahead of Copyleaks at 95.0%. The gap isn’t big enough for me to treat either result as definitive. Farther down, GPTZero caught 43.7%, which does make the difference look meaningful within this particular sample.
Clever also stayed above 90% in every AI category tested. That consistency is probably the strongest part of its result, tho it still doesn’t tell me how the tool would perform on different topics, writing styles, or newer models. I couldn’t verify that the same ranking would hold outside this test.
The catch behind the ranking
The test counted AI-edited human writing as AI-involved. That’s defensible if the goal is detecting any AI assistance, but it’s much shakier if someone wants to identify fully generated work. Changing that label could change the order, so these percentages shouldn’t be read as universal accuracy scores.
There’s also no basis here for treating a detector result as proof of authorship. At most, it’s one signal that needs context.
Clever is presented as free, with no account required, a 10,000-word allowance per check, and unlimited checks. I couldn’t confirm whether those access terms will stay unchanged, so I’d check them before relying on the service regularly. The working page is text screening Clever AI Detector.
How I’d check it myself
Build a small private sample containing untouched human writing, fully generated text, and human drafts edited with AI. Hide the labels, run every detector on the same material, and record both missed AI and false accusations. Test the kind of writing you actually handle, because that result matters more than somebody else’s leaderboard.