Key Takeaways
- An AI detector flagged Stanley Druckenmiller's Wall Street Journal op-ed as AI-generated, and the debate that followed over whether AI-assisted writing is cheating missed the point entirely.
- AI detection tools test whether a machine touched the words, not whether the writing is original, and they get even that narrower question wrong constantly. Non-native English writers take the brunt of it, with false positive rates above 60 percent and one detector misflagging nearly all of their essays.
- AI measurement asks a different question, how much a person added beyond what an AI model would hand anyone by default. Hupmapper scored Druckenmiller's op-ed 211 out of 300 on originality intensity, well clear of the AI baseline, despite the detection flag.
- Hupside calls this standard Original Intelligence, the measurable value added beyond the AI baseline. Hupchecker scores it for people and teams, Hupmapper for a single piece of writing, giving anyone something concrete to measure their work against.
Last week, an AI detection tool called Pangram flagged a Wall Street Journal op-ed about the Treasury’s bond buyback plan by investor Stanley Druckenmiller as AI-generated. Druckenmiller confirmed he had used AI while writing it and said he wasn't embarrassed, comparing it to using a calculator instead of doing long division by hand. The conversation that followed had nothing to do with the bond market. Instead, it turned into a tit-for-tat argument over whether Druckenmiller or a machine had produced the sentences, and almost no one weighed in on whether his opinion on the bond market held up.
I wrote about this story on LinkedIn and argued that the debate should be about whether or not the piece contained an argument a competent AI model would have produced for anyone who typed the same prompt, not if AI was used at all. I ran Druckenmiller's op-ed through a beta version of Hupmapper, and it scored a 211 out of 300 on originality intensity, comfortably clear of the AI baseline. Detection flagged the writing, and measurement found the value in spite of that.
What AI Detectors Look For
AI detection is built around a single question: is this writing human, AI, or some blend of the two? Now that AI involvement is the default rather than the exception, this question is less relevant, and doesn’t even account for how often detectors get it wrong.
AI detectors run a sample of text against the statistical fingerprints common in AI writing and output odds that a machine was behind it. That approach carries two structural problems. The first is that it's a moving target: detectors learn to spot “AI tells,” AI providers learn to obscure them, and the pattern-matching resets every time either side adjusts. The second is who those patterns end up penalizing among writers who never touched a chatbot.
Why Non-Native English Writers Get Flagged More Often
The writing patterns detectors treat as signs of AI overlap heavily with how people who learned English as a second language write. Uniform sentence length, a narrower vocabulary range, and fewer of the small idioms and contractions are all common in AI-generated text. They're also common in writing by people who were taught English's rules explicitly rather than having absorbed them natively. A detector trained to flag consistent, rule-following prose as machine-made will flag both the same way.
AI detection numbers show this pattern frequently and proves a structural bias. The worst-performing detector in a Stanford-led test misflagged 97.8 percent of TOEFL essays by non-native English speakers as AI-generated. Across all seven detectors in that test, the average false positive rate came to 61.3 percent. Hupside ran a similar check on college admissions essays and landed in the same territory, with well over 60 percent of essays by non-native English speakers wrongly flagged as AI-generated.
The stakes here are concrete. A flagged essay can trigger an academic integrity review, and a flagged writing sample can cost a job applicant an interview. Detector makers themselves admit these scores aren't solid ground for a grading decision or an integrity case, and outside researchers studying the tools independently have landed on the same warning, yet the tools keep getting used for exactly that, and are disproportionately used against people who are already writing in a language that isn't their first.
Druckenmiller's case sits on the other side of that same coin. Pangram's flag was accurate. He had said so himself. Knowing a machine was involved never told anyone whether the argument itself went beyond what that machine would have produced for anyone else, which is the more important distinction.
Detection and Measurement Ask Different Questions
Detection and measurement aren't two versions of the same check. They ask different questions entirely, and only one of them tells you anything about the content of the writing itself. I laid out the full side-by-side breakdown, question asked, what each approach can and can't tell you, what each was built for, in AI Measurement vs. AI Detection.
Identification made sense back when AI use was rare enough to be worth flagging on its own. Now that it's the default starting point for most writing, Hupmapper is built to answer what detection can't: how far past that starting point a specific piece of writing goes.
What Originality Intensity Measures
Hupmapper doesn't ask whether AI was involved in a piece of writing, and it doesn't try to guess. It scores how far the writing sits beyond what an AI model would already produce, on a 0 to 300 scale measured against the AI baseline rather than against other writers, and it returns something closer to a digital fingerprint of the work than a probability score. The same measurement underlies Hupchecker, just applied to a person or a team instead of a single piece of writing.
Druckenmiller's op-ed is a clean example. The polished sentences in that piece are common enough that any model could produce them. The thesis, the specific read on the bond market from someone who has spent four decades trading it, is the part no model was going to generate on its own. That's the original signal, and it's what Originality Intensity is built to find, regardless of what tools were used to get the words on the page.
Where This Plays Out Beyond One Op-Ed
This pattern shows up anywhere written work gets evaluated at scale. When most of what a team publishes comes out of the same handful of AI models, one person's work gets harder to tell from another's. Admissions readers face a tighter version of the same problem, evaluating far more applicants per reader than a manager reviews reports from a team. According to Hupside's college admissions research, AI now touches most applications before they're submitted: about half the pool leans on it to brainstorm, and one in five hand it the pen for a full first draft. That's the same reason Hupside built Hupchecker for teams and Hupmapper for individual pieces of writing, rather than one tool trying to cover both.
Why Measurement Should Be the Standard, Not Identification
That same homogenization pattern shows up in the underlying data, not just anecdotally. The scientists behind Hupside, Dan Johnson, PHD, and Dr. Adam Green, ran the numbers on more than 160,000 college admissions essays, combining a lab experiment with two real-world data sets in work that appeared in Information Systems Research. Essays got more varied on the surface as AI assistance rose. Underneath that surface, the ideas got more alike, both inside single essays and across the applicant pool at large.
Hupside refers to this as value signal collapse. When anyone can generate the polished baseline instantly, the baseline stops being worth anything, and the only part of a piece of work that still carries value is what sits beyond it. Novelty and polish, the two qualities detection tools and traditional review both tend to reward, are exactly the qualities AI can now produce on demand for almost anyone. Checking for a machine only ever worked while machine involvement stayed the exception. Measuring how far someone got past it keeps working regardless, which argues for building it into how written work gets evaluated day to day rather than saving it for a special-occasion check.
See Where Your Own Writing Lands
Hupside builds the infrastructure for measuring value beyond the AI baseline, in people, in teams, and now in a single piece of writing. Hupchecker scores that value across the people and teams doing the work today. Hupmapper, live now, scores it one document at a time, the same 0 to 300 Originality Intensity read that put Druckenmiller's op-ed at 211.
For the story that started this one, read my LinkedIn post on the Druckenmiller op-ed. For the fuller research case behind measurement, read AI Measurement vs. AI Detection. Then bring your own writing to Hupmapper and see how far past the baseline it goes.




