Copyleaks AI Detection API Guide: Sandbox and Safe Parsing

Introduction

Connecting an AI detector to a publishing workflow creates two separate problems: making a valid request and deciding what a result is allowed to change. A successful HTTP response solves only the first. This Copyleaks AI Detection API guide focuses on a low-risk integration path: understand the synchronous text endpoint, keep simulated results separate from live observations, and parse changing response fields without turning a missing value into an accusation or an automatic approval.

This is documentation-based research, not a hands-on Copyleaks accuracy review. We checked the linked official documentation on September 28, 2026. We also ran offline unit tests against our own request-builder and parser using synthetic fixtures. Those tests do not contact Copyleaks, consume detection credits, or establish that the service will return a particular result. For actual detector evaluation, use the methodology in our testing guide; for a disputed flag, see our human-writing case-study workflow.

Choose the Text Endpoint Before Building a Queue

The official text-detection guide describes a synchronous call: submitted text and detection results travel through the same request-response cycle. The endpoint is POST https://api.copyleaks.com/v2/writer-detector/{scanId}/check. Do not assume that every Copyleaks product requires a webhook simply because other document-processing routes are asynchronous. Select an endpoint for the actual input format and job you need, then read that endpoint’s contract.

A synchronous vendor endpoint does not require your website to block a customer’s browser while it runs. Our architectural recommendation is to put the request inside a background worker, store a local job record, and expose status to the frontend. That worker queue is your application’s design, not evidence that the vendor endpoint is asynchronous. Keeping the distinction explicit prevents confusing a local job identifier with a vendor callback or treating a network timeout as a completed scan.

For an editorial system, define states before writing the integration: waiting for submission, response received, response invalid, review required, and simulated-only. None of these states should mean “human author proved.” Publication should remain a separate decision based on sources, originality, suitability, and the rest of the editorial record.

Sandbox Is a Wiring Check, Not Free Accuracy Testing

The Copyleaks quickstart distinguishes sandbox: true, which provides simulated results, from real detection that consumes credits. The endpoint reference lists the default sandbox value as false. Explicitly set the field instead of relying on an SDK or server default. Otherwise, a small configuration omission can change both what the response means and whether the request spends credits.

Use three clearly labeled layers:

Layer What it can establish What it cannot establish
Offline fixtures Your code’s behavior for supplied inputs and errors Vendor availability, billing, or detection quality
Vendor sandbox Integration behavior against a simulated vendor response Performance on a genuine human or AI document
Authorized live scan One saved service observation for the submitted input General accuracy or proof of authorship

Our small helper intentionally builds only sandbox requests. It does not include a network client, access token, or a switch that silently turns the request into a paid scan. A team adopting it must add its own authorized transport layer; live use should require an explicit budget and data-handling decision. This separation makes it harder to reuse a demonstration script accidentally as an unattended production scanner.

Authentication and Input Boundaries

The authentication documentation describes exchanging an account email and API key for an access token, then sending that token in a Bearer header. It documents a 48-hour token lifetime. Keep both credentials on the server, avoid logging them, and reuse a valid token rather than logging in for every document. Cache expiry information rather than assuming a token created yesterday must still work.

The text endpoint reference specifies a text length from 255 to 100000 characters and a scan identifier from 3 to 36 characters. Our helper accepts only letters, digits, and hyphens in identifiers as an additional local safeguard, not as a claim that the vendor allows no other characters. It rejects out-of-range text rather than trimming it silently. A truncation changes the document being examined and should be visible in any later audit.

Here is the request shape returned by our helper; this illustration is not an executed API call:

{
  "method": "POST",
  "url": "https://api.copyleaks.com/v2/writer-detector/example-001/check",
  "json": {
    "text": "<your authorized text within the documented length range>",
    "sandbox": true
  }
}

Replace the placeholder with suitable text before an authorized sandbox call. It is deliberately not a detector-ready sample. Do not paste private customer documents into a trial request just to make a demo work. Decide which documents may leave your system, who can access saved inputs, and how long the application should retain them before connecting the transport.

Parse the Response Without Depending on a Deprecated Field

There is a compatibility detail worth checking before copying a quickstart. The response specification identifies section classifications as 1 for human and 2 for AI-generated text. It also marks the per-section probability field deprecated with a planned removal date of July 2026, while documentation examples still show that field. We have not made a live request to resolve the difference. Do not claim that the field has already disappeared everywhere, or make an essential workflow depend on its presence.

Our parser uses section classification, the documented summary values, and modelVersion; it does not require section probability. Missing summary data, unexpected classes, non-finite numbers, or an absent model version produce an error instead of silently becoming a zero AI score. The key policy is that a malformed response must not default to “human.” A valid response is labeled simulated_only when the caller says it came from sandbox and review_required otherwise.

These choices are conservative application rules, not vendor guarantees. Preserve the raw authorized response separately when appropriate, alongside the scan identifier, input version, mode, time, and model version. If you later change parsing behavior, an archived record helps explain why a document’s workflow state changed without pretending the detector’s underlying observation changed too.

Eight Offline Checks and Their Limits

Our Python helper has eight unit tests: sandbox is explicit; text boundaries are enforced; identifiers cannot insert URL paths; missing deprecated probability does not break parsing; a live-mode fixture still requires review; missing summary data fails; invalid numeric values and unknown classes fail; and the mode must be an actual Boolean. The fixtures are constructed test data, not vendor responses. A high synthetic score is used only to exercise code branches, not to measure any product. Here is the checked helper; it intentionally makes no network requests:

import math
import re


def sandbox_request(text, scan_id):
    if not isinstance(text, str) or not 255 <= len(text) <= 100000:
        raise ValueError("text length outside this example's endpoint limits")
    if not isinstance(scan_id, str) or not re.fullmatch(r"[a-zA-Z0-9-]{3,36}", scan_id):
        raise ValueError("use 3-36 letters, digits or hyphens for this example's scan id")
    return {
        "method": "POST",
        "url": f"https://api.copyleaks.com/v2/writer-detector/{scan_id}/check",
        "json": {"text": text, "sandbox": True},
    }


def review_response(payload, *, sandbox):
    if type(sandbox) is not bool:
        raise ValueError("sandbox mode must be explicit")
    if not isinstance(payload, dict) or not isinstance(payload.get("modelVersion"), str):
        raise ValueError("model version missing")
    summary = payload.get("summary")
    results = payload.get("results")
    if not isinstance(summary, dict) or not isinstance(results, list) or not results:
        raise ValueError("summary or sections missing")
    for name in ("human", "ai"):
        value = summary.get(name)
        if type(value) not in (int, float) or not math.isfinite(value) or not 0 <= value <= 1:
            raise ValueError("invalid summary; do not default to zero")
    classes = []
    for row in results:
        if not isinstance(row, dict) or type(row.get("classification")) is not int or row["classification"] not in (1, 2):
            raise ValueError("unknown section classification")
        classes.append(row["classification"])
    return {
        "state": "simulated_only" if sandbox else "review_required",
        "model_version": payload["modelVersion"],
        "section_classes": classes,
        "summary": {"human": summary["human"], "ai": summary["ai"]},
    }

Passing these tests shows that the checked version of our helper follows those local rules. It does not show that authentication works on your account, that your plan permits a scan, that the response schema will never change, or that Copyleaks can classify your documents correctly. The distinction matters because an integration tutorial can offer useful, reproducible evidence about its own code without presenting a mocked result as an experiment on the service.

For your own application, add tests for timeouts, expired tokens, interrupted jobs, and schema changes before enabling a transport. Run those failure tests against mocks first. If you later conduct live evaluation, store its inputs and outputs in a separate evidence collection with a clearly stated method and limitations; do not merge it with unit-test fixtures.

Errors, Rate Limits, and the Budget Boundary

The error reference separates request, authentication, payment, and service errors. Treat these as different recovery paths. An invalid request needs correction, an authentication failure needs credential or token handling, and insufficient credits need a budget decision. None is a reason to turn on unlimited retries. Keep an unresolved job visible instead of marking a document processed because an exception was caught.

The rate-limit guide documents a default account limit of 10 requests per second, with a stricter login limit of 12 requests per 15 minutes, and recommends backoff for 429 responses. A team should still set a lower local concurrency and a bounded retry count appropriate to its needs. A timeout is not proof that a remote operation did not occur; do not invent an idempotency guarantee from the presence of a scan ID. Investigate ambiguous outcomes before resubmitting paid work.

Separate limits on attempts, concurrency, and expenditure. An attempt cap alone cannot guarantee a dollar budget if pricing or credit consumption changes. Read the account’s current plan and official billing information before authorizing real detection. Check the official website for the latest pricing. This guide quotes no plan price and enables no paid API calls.

Final Verdict

Start with an explicitly labeled sandbox workflow and offline tests for your own parser. Keep credentials out of the browser, reject invalid input visibly, and avoid relying on a response field the vendor marks deprecated. When you authorize live scanning, preserve enough context to audit the result and route it to a review process rather than an automatic authorship verdict. The useful purchase decision is whether the documented interface fits your workflow and budget—not whether a simulated demonstration proves accuracy.

Official Sources and Disclosure

Sources checked September 28, 2026: text guide, quickstart, endpoint, authentication, response fields, errors, and rate limits. Product behavior described here comes from those documents, not a paid scan. AI assisted with drafting and organization; the offline checks apply only to our own code. This article may contain affiliate links. Payment or sponsorship does not determine its conclusions.

AscendProse Intelligence

Stay Ahead with Verified AI Software Insights

Join 1,000+ creators, marketers, and SaaS builders receiving monthly benchmark research, exclusive promo deals, and practical AI workflow breakdowns. Unsubscribe at any time.

Limited-Time Offer Ready to try Copyleaks?
Get Started Free →