AI Detectors: What They Are, How They Work, and Whether They Really Work to Detect ChatGPT Text

Dan Patiño

AI Strategy & Innovation at Coderhouse

Artificial Intelligence

AI Detectors: What They Are, How They Work, and Whether They Really Work to Detect ChatGPT Text

Published on

AI detectors are tools that try to estimate whether a text was written by a person or generated by a model like ChatGPT. They analyze statistical patterns of the language, but they aren't infallible: they have significant error rates and can flag legitimate human texts as "AI". They serve as an indicative signal, never as definitive proof.

With the explosion of generative AI, teachers, editors, and content teams share the same question: how to know if a person wrote this? Then GPTZero, Originality.ai, Copyleaks, and dozens of competitors appeared. But do they really work? Here we analyze how they operate, how reliable they are, and when it makes sense to use them.

How AI detectors work

These tools rely on two main metrics of the text.

Perplexity

It measures how predictable a text is. AI models tend to choose the most probable words, generating texts of low perplexity, that is, very "predictable". A human is usually more unpredictable in their choice of words.

Burstiness

It measures the variation in the length and structure of the sentences. People naturally alternate short and long sentences; AI tends toward a more uniform rhythm. Detectors look for that "monotony" as a telltale signal.

The problem is that these signals are probabilistic, not certainties. A very polished human text can seem "AI", and a well-edited AI text can pass as human.

Comparison of the best-known tools

Tool

Approach

Typical use

GPTZero

Perplexity and burstiness, analysis per sentence

Academic and educational field

Originality.ai

AI detection + plagiarism, SEO-oriented

Editors and content agencies

Copyleaks

Multilingual AI detection + plagiarism

Companies and institutions

None publishes a 100% accuracy, and the figures they promote are usually measured in laboratory conditions, not in real use with edited or translated text.

The problem of false positives

This is the critical point. A false positive (flagging a human text as AI) can have serious consequences: unfair accusations of cheating against students, content rejected for no real reason. Cases have been documented of texts written by non-native English speakers erroneously flagged as generated by AI, a worrying bias.

OpenAI itself discontinued its AI text classifier due to its low hit rate, publicly acknowledging how difficult this task is. If the creator of ChatGPT didn't achieve a reliable detector, it's a good idea to take others' promises with a grain of salt.

Media like The Verge covered these limitations extensively and the risks of blindly trusting these tools to make important decisions.

So, do they work or not?

They work as one more signal, never as a verdict. They can help an editor review a suspicious text with more attention, but they shouldn't be the only basis for accusing someone or rejecting a piece of work. Human judgment remains irreplaceable.

Instead of obsessing over detecting, many teams choose to integrate AI transparently into their processes. Understanding how to use these tools well is more productive: for example, applying AI in tasks like product discovery to research users and validate hypotheses.

Learn to use AI in your favor with Coderhouse

Instead of fearing AI, master it. These courses help you integrate it with professional judgment:

Frequently asked questions

Are AI detectors 100% reliable?

No. None reaches total accuracy. They work with probabilities and can be wrong both flagging human text as AI (false positive) and letting AI text pass as human (false negative).

How does one of these tools detect a text generated by ChatGPT?

They analyze statistical patterns like perplexity (how predictable the text is) and burstiness (the variation in the sentences). AI tends to be more predictable and uniform, and that's what these tools look for.

Why did OpenAI close its own AI detector?

Because its hit rate was low and it generated too many errors. It was a public acknowledgment of how difficult it is to reliably distinguish a human text from one generated by AI.

Can I use an AI detector to accuse someone of cheating?

It's not recommended as sole proof. Given the high risk of false positives, a positive result should be just a signal to review with more attention, never a definitive accusation.

About the author

Dan Patiño

I'm Dan Patiño, head of AI Strategy & Innovation at Coderhouse. My day-to-day work involves merging the tactical management of e-commerce (CRO, Email Marketing and SEO) with the development of disruptive solutions. I specialize in building internal AI-powered apps to automate tasks and boost innovation within the team. I firmly believe that technology is strategy's best ally. To dive deeper into my professional journey, I'll be waiting for you on my LinkedIn profile.

English

© 2026 Coderhouse. All rights reserved.

English

© 2026 Coderhouse. All rights reserved.

English

© 2026 Coderhouse. All rights reserved.

English

© 2026 Coderhouse. All rights reserved.