Blog

Why a Benchmark Would Deliberately Hide Its Own Results From the Public

A benchmark publishes scores, methodology, and raw data, then withholds the one thing that seems most obviously useful to share: the actual rewritten text each tool produced. That choice looks like it undermines the whole point of publishing a benchmark in the first place. It is actually the opposite, a deliberate safeguard against a specific, well-documented way benchmarks like this one get quietly undermined over time.

Understanding the reasoning behind this choice changes how the entire benchmark should be read, from an incomplete disclosure into a carefully designed tradeoff between openness and long-term usefulness, one that applies to every tool tested, not only the one running the test.

The Specific Risk a Humanizer Benchmark Creates

An AI humanizer benchmark exists to measure how well a rewriting tool helps text pass as human-written against a detector. The output of that benchmark is, by definition, a large set of text specifically optimized to defeat detection. Publishing that text openly hands detector developers a ready-made training set of exactly the content their own models most need to learn to catch.

This is not a hypothetical concern. Pangram’s own published research, a paper on arXiv describing a method called DAMAGE, documents training a detector to resist evasion more effectively specifically by including humanizer output in its training data through data augmentation. A public benchmark archive of humanized text is, in effect, free training material for the next generation of the very detectors the benchmark is measuring performance against.

This creates a strange incentive problem for anyone publishing this kind of comparison. The more useful and comprehensive a public archive of humanized text becomes, the more directly it helps the opposing side of the exact contest the benchmark was designed to measure, which is precisely the tension a withholding policy is built to defuse.

How Phrasly’s September 2026 Benchmark Handled This

Phrasly’s humanizer benchmark, comparing its own Ultra model against StealthGPT, WriteHuman, and Undetectable AI across 72 AI-generated texts, withholds every output from every tool, including its own. The public download for each tool lists every input, every Pangram score, every hallucination count, and a SHA-256 fingerprint of each output, but not the output text itself.

That fingerprint is the mechanism that makes withholding compatible with verification. A SHA-256 hash is a short string generated from a piece of text that changes completely if even one character of the original is altered, which means anyone who later receives the actual withheld text can compute its hash and confirm it matches the one published in the CSV, proving the text they received is exactly the one that was scored, without the text itself ever needing to sit in a public archive.

This approach also applies the same rule evenhandedly to every tool in the comparison, including the one built by the company running the test. Phrasly’s own Ultra outputs are withheld under the identical policy applied to StealthGPT, WriteHuman, and Undetectable AI, rather than carving out an exception for its own product while restricting access to everyone else’s.

Who Can Actually Request the Withheld Texts

The benchmark is not permanently sealed. Researchers, journalists, and the specific providers tested can request the withheld outputs directly, describing their intended use. This keeps a path open for genuine scrutiny without defaulting to the kind of unrestricted public release that would create the exact training-data problem the policy exists to avoid.

Access to the withheld texts comes with specific conditions:

  • Requesters must agree not to use the texts to train or tune AI detectors
  • Requesters must agree not to republish the texts publicly
  • Every request is reviewed individually, with a response promised within 10 business days
  • Nothing is released automatically, regardless of who is asking

This structure lets the Phrasly Benchmark stay genuinely checkable by the people with a real reason to verify it, researchers studying detection methods, journalists reporting on the results, and the competing tools being compared, without the tradeoff of making the same material freely available to anyone who might use it to retrain a detector instead.

What This Practice Says About Benchmark Design Generally

Most published comparisons in this space come from the companies selling the tools being compared, and most never address this contamination risk at all, either because it was not considered or because the convenience of a fully open archive outweighed the concern. A benchmark that explicitly accounts for this risk, and builds a structured, conditional access process around it, is making a methodological choice that costs the publisher convenience and public-facing completeness in exchange for not actively degrading the exact detection technology the benchmark depends on to mean anything.

This also means the absence of a visible output archive should not, by itself, be read as evidence of lower quality. A benchmark that withholds its outputs for this specific, documented reason has made a more deliberate methodological choice than one that simply never considered the question.

Transparency and Restriction Are Not Actually in Conflict

It is easy to assume more openness is always better when evaluating how trustworthy a benchmark is. This case shows why that assumption does not always hold. Withholding the most sensitive piece of data while publishing everything else needed to verify it, scores, methodology, and cryptographic fingerprints, can be the more responsible design, not a lesser one, once the actual risk that openness would create is taken seriously.

The more useful question to ask of any benchmark is not whether every piece of underlying data was published, but whether enough was published, in the right form, for an independent party to actually check the claims being made. A fingerprint and a conditional access process can satisfy that standard just as well as a fully open archive, sometimes better, once the specific risks involved are properly accounted for.

For more on how AI detection and humanization technology actually interact, further reading on the Phrasly blog covers the underlying research for anyone evaluating a benchmark’s design choices.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button