When a Full Stop Becomes AI: Questioning the Reliability of AI Detection Tools
Generative AI is becoming the foundation of more content, leaving many questioning the reliability of AI detectors.
I recently ran a simple experiment, partly for curiosity, partly to see how absurd it could get.
I asked four leading AI detection platforms a seemingly absurd question: “Was this single full stop (aka a period) generated by AI?” To my surprise, all four platforms produced the same answer: “100% AI-generated”. One full stop. What?
This experiment may seem trivial to most, but it reveals a far more serious concern: the accuracy and consequences of AI detection tools used in academia, publishing, and corporate environments.
How AI Detectors Work
AI detection platforms do not evaluate text by meaning. They analyse patterns across massive datasets: novels, academic papers, news articles, white papers, and all other printed or digital material. They identify statistical cues in sentence structure, word choice, and syntax, comparing them to learned patterns.
These tools detect patterns, not intent.
My full stop example demonstrates how even the smallest unit of writing can be misclassified, showing that the technology is far from infallible.
But why does this matter?
Well, the implications are significant. Universities increasingly deploy AI detectors to flag student work. Publishers consider them for manuscript assessment. Corporations use them to evaluate content for compliance or branding.
AI detection tools are now becoming a standard part of academic oversight. Some tertiary institutions even integrate these detectors directly into their online submission platforms. These systems display the original submission alongside all changes made post-submission, ostensibly to ensure originality and integrity. So, students can’t adjust their work in real time? Absurd, right?
A real-world example illustrates the stakes. Recently on TikTok, a student under the name “Sawyer” recounts a lecture experience that has since garnered over 23 million views, 3 million likes and close to 10k comments, most of which view the professor’s stance as clearly misinformed.
Here’s the story.
At the start of class, the professor places an essay on the screen and asks the class who’s work it is. Sawyer raises his hand but is unsure why it has been singled out.
The professor presses, asking, “Did you write this?” Sawyer replies confidently, “Yes, of course I did.” The professor responds, “Ok, see me after class,” leaving Sawyer to endure the weight of the interaction for the entire session.
After the lecture, Sawyer approaches the professor again and reiterates, “Hi, that’s my essay.” The professor points to a solitary em dash in the work and asks, “What’s this?” Sawyer, confused, answers, “What do you mean, it’s an em dash.” The professor insists, “I think I know what an em dash is. Don’t play smart with me. That means you used AI in your assignment.”
Sawyer doubles down, insisting he wrote the essay entirely himself. The professor counters, “Well, I’ve only seen this in AI papers.” Sawyer suggests running the essay through GPT-Zero¹, which returns “100% generated by the author.” The professor dismisses the result, saying, “You think I believe in that cute computer stuff?”
The encounter ends with the professor telling Sawyer to leave the class immediately and that he is failing the subject. Sawyer is then required to submit a written explanation to the department head to avoid failing.
This viral TikTok illustrates a much broader issue: students are increasingly being judged by AI detection results, sometimes in public and without clear justification. The technology, while positioned as a safeguard, clearly creates stress, confusion, and potentially unfair academic consequences.
One TikTok commentor quipped: “Oh sure, let’s fail students for using correct punctuation now. Em dash = AI? Got it. Next up: semicolons get you expelled and commas trigger investigations. Let’s just allow a broken education system to rewrite the English language (smh)”
But academia is doubling down.
As flagged above, tertiary institutions now use these detectors not only to flag suspected AI use but also to monitor changes made post-submission. While designed to maintain academic integrity, such systems amplify anxiety, particularly when instructors place disproportionate trust in algorithmic outputs over direct context or discussion with the author.
If a detector can confidently classify a single punctuation mark or em dash as AI-generated, what confidence should decision-makers place in broader applications like student essays, research reports, or internal communications?
Decisions made under the assumption of accuracy can misclassify text and create unjust outcomes.
Additional Simple Examples
My simple full stop experiment is only the beginning. Other recurring misclassifications include:
Single commas in complex sentences flagged as AI-generated.
Short phrases such as “In conclusion” or “As discussed above” receiving high AI scores.
Common words appearing in unusual combinations, like “and,” “the,” or “but,” flagged incorrectly.
Short-form social media posts or tweets misclassified despite being entirely manually composed.
These examples, drawn from crowdsourced experiences shared online, underscore a critical point: AI detection tools are pattern-based, not context-aware. Errors are not anomalies, they’re expected outcomes.
This, therefore, raises urgent questions for decision-makers:
Academia: Are students being unfairly penalised? Are instructors placing undue reliance on statistical outputs rather than evaluating text in context?
Publishing: Could manuscripts be rejected based on statistical quirks rather than content quality?
Corporate Communication: Are internal memos, reports, or client comms misclassified, potentially creating legal or reputational risk?
Reliance on flawed AI detection tools risks distorting accuracy, fairness, and trust across multiple domains.
The Pandora’s Box of AI Detection
AI detectors, by design, cannot interpret nuance, style, or intent. They can only measure deviation from learned patterns. My “full stop experiment” demonstrates how even the most minimal text element can trigger false positives².
AI detection is clearly a Pandora’s Box.
Yet organisations continue to adopt these tools, assuming accuracy and infallibility. The consequences of misinformed decisions are profound: academic penalties, rejected manuscripts, misaligned corporate communications, and broader erosion of trust in oversight systems.
So, if a single period can be misclassified with 100% certainty, how much confidence should be placed in AI detectors for more substantive text? The answer is not “none,” but it’s certainly severely limited. Critical oversight, careful evaluation, and awareness of limitations are essential.
AI detection results must be treated as advisory at best, not determinative.
Until these systems evolve to interpret context and meaning more accurately, the risk of misclassification will persist. Failing to question these tools now risks embedding error and distrust into institutions that are meant to shape knowledge, creativity, and corporate decision-making.
One practical insight for spotting AI-generated text: look for stylistic deviations. If a writer has never used em dashes — or other distinct punctuation — or if their latest post or essay significantly deviates from their usual style, there’s a strong chance an LLM was involved. Scour your own network, from colleagues to LinkedIn followers, and you’ll quickly notice patterns emerge.
¹GPT-Zero is an AI detection tool designed to determine whether a piece of text was likely written by a Large Language Model (LLM) like ChatGPT, Gemini, DeepSeek or similar platforms. It was developed by educators and computer scientists to help detect AI-generated content in academic settings. The tool does not “read” or understand meaning. It assesses statistical patterns and outputs a probability that the text was AI-generated.
²A false positive happens when an AI detector flags a piece of text as AI-generated even though it was written without AI assistance. In this example, if GPT-Zero labels a student’s manually written essay as “100% AI-generated,” that’s a false positive. It’s essentially a “false alarm”. The system signals a problem where there is, in fact, none.
If you liked reading this piece, you may also be interested in my June 2025 article titled “Is the Em Dash Being Bullied — Or Are Writers Just Tired of Being Misread?”




