AI has made polished writing and analysis cheap and abundant, so the old ways of checking whether someone did the work no longer work. The instinct to test for original thinking is to reach for an AI detector, but that tool answers the wrong question. AI detection tries to determine whether a machine touched a piece of text, while AI measurement asks how much value a person added beyond what AI alone would have produced. Treating those as the same question makes organizations miss the signal they need.
Key Takeaways
- AI detection tools try to answer whether AI touched a piece of writing, but research has found wide swings in accuracy along with bias against certain writers.
- AI measurement asks how much value a person added beyond what AI alone would have produced. This is the approach organizations evaluating written work should be adopting as the default, not a niche alternative to detection.
- Employers and admissions offices are both finding that the old signals of quality no longer separate people because AI can produce a polished draft for almost anyone.
- Original Intelligence provides a name and a standard. Hupchecker measures it across people and teams, and Hupmapper applies that same measurement to individual pieces of writing, so anyone can check their own work against it directly.
What AI Detection Tools Measure
AI detection tools score text against patterns common in AI-generated language and return a probability that a machine produced it. That probability is far less reliable than it sounds. A Stanford-led study published in the journal Patterns tested seven widely used detectors against TOEFL essays written by non-native English speakers and found an average false positive rate of 61.3 percent.
Hupside's own research into college admissions cites a similar rate, finding that more than six in ten essays from non-native English speakers get wrongly flagged as AI-generated. That figure tracks closely enough with the Stanford findings above that it is worth naming plainly. This is a well-documented problem, and it’s showing up wherever anyone measures for it.
This means that an entire category of tools built to answer "did AI write this" can misfire based on a writer's language background or sentence structure, sometimes more often than it gets the answer right. Even the companies that build these detectors acknowledge they aren’t reliable enough for high-stakes decisions like grading or academic integrity enforcement, and independent researchers have reached the same conclusion.
Why Detection Is the Wrong Question
AI is the floor nearly everyone starts from now, and is part of why Hupside treats AI as a neutral baseline rather than something to catch people using. The useful question is how far past the AI floor they went, which is a question detection was never built to answer. Identification made sense when AI use itself was rare enough to be a meaningful signal, but measurement is now what organizations should be evaluating instead.
What AI Measurement Gives an Employer
Employers are running into this detection vs measurement problem at enterprise scale, well beyond any single essay or resume. As more businesses use AI to assist them, more and more businesses are reporting negative consequences tied to AI inaccuracy. When most of a team's output runs through the same handful of models, individual contributions start to look alike on the surface, whether the person behind them is meaningfully improving on that starting point or simply forwarding it along.
This over-use of AI leads to low-value, derivative output that passes every polish check but adds nothing distinct on its own, aka AI Slop. Hupside uses Hupchecker to combat this issue by assessing an individual's Original Intelligence and providing an OIQ score and an OIQ archetype for each team member. This provides insight into how your team members interact with AI, and because the measurement lives with the person, it holds up across outputs, and it keeps working even as the underlying AI models change.
What AI Measurement Gives a College Admissions Office
College admissions offices feel a version of this pressure too, with far less reader capacity per applicant than an employer has per team member. Roughly half of applicants now use AI to brainstorm their essays, and about one in five use it to generate a first draft, according to college admissions research. That means the essay, once one of the more reliable signals of independent thinking in an application, is carrying a lot less information than it used to.
Across roughly 2,200 essays, human-written submissions collectively generated up to eight times more novel ideas than AI-assisted ones, once ideas were scored for how far they moved beyond an AI-typical baseline rather than for style or polish. Research led by Georgetown neuroscientist Dr. Adam Green, a Hupside co-founder, found that originality-based scoring identified original thinkers more accurately than trained admissions readers working from the essay alone.
Why Measurement, not Identification, Should Be the Standard
The pattern holds outside any one hiring pipeline or admissions cycle. Georgetown's Dr. Adam Green and creativity researcher Dan Johnson combined a controlled experiment with two large natural experiments across more than 160,000 U.S. college admissions essays, published in Information Systems Research, and found a striking homogenization effect. AI-assisted essays used more varied vocabulary than unassisted ones, but the ideas underneath that vocabulary grew more repetitive, both within individual essays and across the applicant pool as a whole. AI’s ability to produce polished and novel content on demand is leading to signal collapse; meaning the old signals of capability stop separating people now that AI can generate these signals easily.
Identification was always going to run out of runway once AI could pass as anyone in any industry. Measurement is built to keep working after that point, which is why it belongs in the standard evaluation process for written work, not on the sidelines as a specialty check.
See How Far Beyond the Baseline You've Gone
Hupside is the Original Intelligence Infrastructure company, built to measure the capacity to create value beyond the AI baseline, and its tools put that same science-backed standard to work. Hupchecker measures Original Intelligence in people and teams today, scoring OIQ across a group of contributors so a manager can see where original value is being generated. Hupmapper brings that same measurement to a single piece of writing, scoring originality intensity so a reviewer can see how much value a person added beyond what AI alone would have produced. Between the two, Hupside gives employers and admissions teams a science-backed measurement of how much of a piece of work belongs to the person who made it that identification alone can’t provide.
Anyone curious how much of their own thinking shows up in a piece of writing, beyond the AI baseline, can run it through Hupmapper directly or learn more about the science behind Original Intelligence at hupside.com.




