When a high score appears, the reject button comes into view first
You put a draft from an outside writer or applicant into an AI detector and get “98% likelihood of AI generation.” The deadline is approaching, and it is unclear what to ask the writer. Turning that score directly into a verdict may be fast, but it is not safe.
Results from AI detectors, including Pangram, should be used not as evidence to judge an author, but as screening signals that indicate which passages need closer review. That is because detectors analyze patterns in finished text rather than the process through which the sentences were created.
There are extreme examples, too. Fritz’s 2026 comparison article reports that GPTZero labeled the U.S. Constitution as AI-written and Originality.ai labeled 『The Da Vinci Code』 as 100% AI-generated. However, the article does not include the full input, run date, product version, threshold, or result logs, so it cannot be treated as an independently reproducible benchmark. It is more accurate to read it as a cautionary example that even well-known human-authored works can be flagged.
What Pangram changed is training data, not proof of authorship
In July 2026, Pangram announced Pangram 4 and an image-detection model after raising $9 million in a round led by Menlo Ventures. Haystack, ScOp, Script Capital, and Cadenza also participated. Funding, however, is separate from independent validation of product performance.
Pangram 4’s core methods are synthetic mirror and hard-negative mining. First, it creates AI documents similar to each human document in topic, length, and style, then uses them as training pairs. It then feeds back cases where human writing was incorrectly classified as AI into the training data to reduce false positives.
There is one important boundary here. This approach learns the statistical differences between human and AI documents more effectively. It is not a technology that proves an author by checking drafts, keystrokes, paste history, or watermarks.
The product divides documents into Human Written, AI-Assisted, and AI-Generated segments, returning a score, confidence level, and character location for each. According to the model card, it analyzes documents with overlapping 512-token windows and then refines results at the sentence level. Adjacent labels may be merged to a minimum length of roughly two sentences, so the displayed colors should not be interpreted as word-by-word authorship origins.
A 0.0041% false-positive rate is a strong result, not a universal law
The figures Pangram has published are useful because their denominators can be checked. But they are all company-reported results, measured by Pangram researchers on predefined evaluation corpora.
| Evaluation target | Published result | Limitation when interpreting it |
|---|---|---|
| English human documents | 41 false positives out of 1 million FineWeb documents, 0.0041% | False-positive rates varied by domain: 0.019% for academic writing, 0.078% for poetry, and 0.049% for recipes |
| English AI-generated text | 1,766 misses out of 519,993 samples from 26 models; 0.3396% false-negative rate | Math, coding, and multiple-choice prompts were excluded from the evaluation |
| Korean evaluation | 0% false-positive rate for pre-2022 FineWeb2 human documents; 0.5805% false-negative rate for AI-generated text | The Korean-specific sample size was not disclosed, and AI-generated text was evaluated internally using translated and mirrored English documents |
The reported 95% confidence interval for the English human-document false-positive rate was 0.0032%–0.0050%, while the interval for the English AI-generated-text false-negative rate was 0.3242%–0.3558%. The model card also reports a 0.5805% false-negative rate for Korean AI-generated text, but does not provide a Korean-specific sample size. Moreover, this figure came from an internal synthetic evaluation built by translating and mirroring English documents, so it does not represent accuracy across real Korean applications, short ad copy, or specialized drafts. The fact that the official site explicitly supports Korean should be distinguished from a claim that Korean real-world accuracy has been proven.
Use a free account to test the review workflow, not to make judgments
Pangram’s current free plan provides text checks of up to 2,000 words per day without a payment method. You can paste text or upload PDF, DOCX, and RTF files, with no separate installation required. Rather than submitting real materials from the start, first check the results screen and recordkeeping process with two internal validation drafts whose AI-use status you know. If keeping both drafts within a combined 2,000 words is difficult, test one per day.
- Start by confirming upload authorization and length.
Prepare one “original text written directly by a person” and one “human-written draft lightly edited with AI,” with personal and confidential information removed. To test them on the same day, keep their combined length within 2,000 words; otherwise, split them across dates. Do not proceed if authorization to upload to an external service is unclear. - Log in with a free account at the official Pangram page and select Text Detection.
Paste the full draft or upload a supported file to run the check. Pangram 4 specifies a minimum input of 50 words of prose in complete sentences. - Review the flagged segments rather than only the overall score.
Record Human Written, AI-Assisted, and AI-Generated labels, along with confidence levels and locations for each segment. If you connect the API, you can also retain segment scores and character locations from the response. - Compare flagged segments against separate evidence.
Check drafts and version history, quoted originals, source inconsistencies, and abrupt stylistic changes from prior work. Keep scores and independent evidence separate in your records. - Ask the author only about passages that require an explanation.
Give them an opportunity to explain the order in which they wrote and edited the passage and to provide earlier drafts. Also record the test date, model version, input scope, segment results, and basis for the final decision.
The success criterion for the first trial is not getting both drafts perfectly right. You have established the first workflow if the team can record detection scores separately from other evidence and agree on which results trigger additional human review.
One independent piece of evidence is better than two detectors
Running two different detectors and getting the same result does not prove authorship. If the two products rely on similar textual features, they can be wrong together on the same text. A peer-reviewed study comparing 14 detectors also found large performance differences by tool and document variation, and concluded that their results were unsuitable as evidence of academic misconduct. However, that study used 54 English samples from 2023 and did not include Pangram 4.
The University of Glasgow’s 2026 guidance advises examining the full body of evidence—such as earlier drafts, the author’s statement, and nonexistent or mismatched sources—rather than treating AI detection scores as grounds for investigation. The same principle applies to content operations.
You can divide review standards this way.
- If the goal is editorial prioritization, use detection results to choose passages a person should review first.
- For high-impact decisions such as rejection, rescinding an offer, or discipline, do not use a score alone without drafts, sources, and the author’s explanation.
- If there are no standards for upload authorization or personal-data handling, establish internal procedures before using a detector.
- If Korean, specialized fields, or short documents are the main inputs, create an internal baseline from your actual document distribution before defining the scope of use.
Pangram 4’s progress lies in providing more specific evaluation figures and mixed-authorship segments. Still, the question a detector can answer only goes as far as, “What patterns does this text resemble?” Answering “Who wrote it, and through what process?” requires drafts, version history, sources, and the author’s explanation.
If you want to dig deeper
Pangram 4 Technical Report A technical report for directly reviewing synthetic mirror, hard-negative mining, and the makeup of the evaluation corpora. arxiv.org
Pangram 4 Model Card Useful for checking English and Korean evaluation results, error rates, input requirements, and how mixed-authorship segments are handled. pangram.com
GenAI Guidance 2026 - Identifying Potential Concerns University of Glasgow guidance showing which evidence—such as drafts, statements, and sources—to review instead of relying on detection scores. gla.ac.uk


