Introduction
An AI detector flags a document that its author says they wrote. The immediate question is not “Which detector should we trust?” It is “What evidence would justify a decision about this document?” A displayed score can start an inquiry, but it cannot identify an author or replace a writing history. In a saved AscendProse experiment, two texts with documented human origins received high displays from one detector and very low displays from two others. This guide turns those cases into a practical response procedure for editors, educators, and teams. It is not a detector ranking or a claim about how often false positives occur.
The underlying run was completed on September 2, 2026. We submitted the same unmodified input to each of three public interfaces, saved the outputs and screenshots, and checked the human excerpts against their source texts. You can inspect the raw observations and method and use our separate guide to designing an AI detector test if you want to reproduce the process. Here, the useful question is narrower: what should a decision-maker do when a score conflicts with provenance?
Case Study: Two Human Texts, Conflicting Displays
The first excerpt comes from “A Landholder, I.” in Essays on the Constitution of the United States, a collection of writings from the 1787–1788 ratification debate. The second comes from Henry Hunt Snelling’s nineteenth-century photography manual, The History and Practice of the Art of Photography. The Project Gutenberg record for the essays and the record for Snelling’s manual document the source works. Our archived sample files and input hashes identify the exact excerpts submitted; the book landing pages alone are not a substitute for those saved inputs.
sample-id |
Source genre | Sapling AI Detector | Undetectable AI Detector | ZeroGPT AI Detector | Documented human source |
|---|---|---|---|---|---|
human-editorial-02 |
Historical editorial | 100.0% | 1.0% | 0.0% | Constitution essays |
human-technical-01 |
Historical technical writing | 95.6% | 1.0% | 0.0% | Photography manual |
These are the percentages displayed by the interfaces for these exact inputs on the saved test date, not calibrated probabilities that the authors used AI. The cases were chosen after examining the outputs because they illustrated disagreement. We examined six historical human excerpts in that run, not a representative sample of present-day writing. The table therefore does not establish a general false-positive rate, a product ranking, or the chance that any new document will be misclassified.
There is also a time boundary. The test occurred on September 2, 2026. Sapling’s detector page now lists a later September 17 model update. We did not rerun these excerpts on that later version. A reader should not assume that the current interface would reproduce the historical displays. The vendor page itself cautions against using its detector as a standalone determination; the archive lets readers examine what was actually seen in our earlier run.
What to Do When a Human Author Is Flagged
1. Freeze the exact input before debating the output
Save the submitted text, its file or document version, the detector’s full result, the date, and any visible product or model version. Record whether you pasted text, uploaded a file, or changed the sample to fit an input limit. A later rerun on a revised or truncated document answers a different question. If the document contains confidential material, do not upload it to additional public tools merely to settle a dispute; review the service’s data terms and your organization’s rules first.
2. Separate primary provenance from classifier opinion
A draft history, version-control trail, contemporaneous notes, source files, or a historical publication record can speak directly to how a document came to exist. A detector score is an inference from the text presented to a model. These are different kinds of evidence. In our two examples, the preserved excerpts can be traced to historical works; the later interface outputs cannot rewrite that origin. For a modern author, ask for material they can reasonably provide without treating a missing draft history as proof of misconduct. Some legitimate writing processes leave fewer records than others.
3. Investigate disagreement without taking a vote
Running the same text through another detector may reveal that outputs are inconsistent. In the two archived cases, the three interfaces disagreed sharply. That observation should lower confidence in a single automated accusation, but it does not prove which detector is generally correct. Detectors may use different models, input handling, thresholds, and display conventions. Do not average their percentages, count “two against one” as a verdict, or interpret a low display from one tool as a certificate of human authorship. If you compare tools, keep the exact input and conditions documented.
4. Ask a question the author can answer
Tell the author what was flagged, what text was submitted, and what decision is being considered. Invite an explanation and relevant process evidence before drawing a conclusion. For an editorial assignment, that might mean drafts, notes, citations, tracked changes, or a conversation about specific claims in the document. For a historical document, it may mean locating the original publication and matching the excerpt. The goal is to resolve a factual question fairly, not to force an author to defend a secret threshold they cannot inspect.
5. Make the decision from the whole record
Use a simple evidence ladder. First examine direct provenance and the actual document history. Next review the text and sources for the concern that prompted the check. Then use detector outputs only as supplementary signals, with the input, date, and product recorded. If provenance establishes human origin, a high detector display is a misleading flag for that case. If provenance remains uncertain, describe the uncertainty and follow an established review policy; do not quietly convert a score into a finding of AI authorship. Where academic, hiring, or disciplinary consequences are possible, a single automated display is especially inadequate.
6. Record an outcome that can be audited
Keep a short case note: the exact input identifier, source or author-provided records, each tool and date, relevant limitations, the person who reviewed it, and the reason for the decision. This is more useful than a screenshot of one number. It also helps a team notice recurring problems with its process without claiming a statistical error rate from a handful of disputed cases. Delete or retain sensitive text according to the applicable policy, not simply because a detector offers a convenient history feature.
What These Cases Do—and Do Not—Add
The distinctive value of this small study is the combination of traceable human-source excerpts, identical inputs across three public interfaces, saved outputs, and a visible disagreement. It provides a concrete reason to build an appeal and verification path into any detector workflow. It does not tell us why Sapling displayed a high value for either excerpt. Formal historical prose, document length, model training data, and interface changes are all conceivable factors, but this experiment did not isolate them. Attributing the result to one writing style would be an unsupported causal story.
Nor does a source work’s age automatically validate every copied passage. Our evidence chain is the specific excerpt, its saved input hash, the source marker, and the archived detector output. Someone applying this procedure to a new case should preserve that chain rather than assume that a link to a book or a draft folder proves an arbitrary submitted text. The downloadable dataset exposes the recorded numbers, while the testing guide explains how to build a comparable new run.
Limitations
The human texts in this run are public-domain historical English, not a sample of contemporary student work, job applications, newsrooms, or other languages. Only six human excerpts and six known AI inputs were used across three tools. The two cases above were selected because they were notable, so they cannot estimate population performance. The interfaces may have changed after the recorded date; Sapling publicly lists a subsequent update. Different displayed percentages may not share the same calibration or meaning. The source records show the works’ historical origin, while the saved research archive identifies the exact excerpts. No conclusion here should be used alone for academic punishment, hiring, or disciplinary action.
Sources and Method
The public benchmark page contains the run methodology, limitations, and raw output table. The same-input check used the Sapling AI Detector, Undetectable AI Detector, and ZeroGPT AI Detector public interfaces. Human-source provenance comes from the Constitution essays and photography manual records, with exact excerpts and hashes preserved in the experiment archive. The product pages were reviewed again on September 27, 2026; this later source check is not a new detector test. AI assisted with drafting and organization, while the displayed values and source records came from the saved experiment.
Final Verdict
If a detector flags writing with credible human provenance, pause the accusation and investigate the record. Preserve the exact input, confirm its origin, ask the author for reasonable process evidence, and document any cross-tool disagreement without treating it as a vote. Our two September 2 cases show why that workflow matters, but they do not predict how any detector will handle a new document today. A fair decision requires more than a number on a screen.