Skip to main content

AI Detector API

Sapling's AI detector API computes the probability that a piece of text is AI-generated (sometimes called "AI slop"), as well as the probability that each constituent sentence and token is AI-generated. Use this AI detection API to flag machine-generated content in user submissions, marketplace listings, reviews, job applications, student work, and content pipelines — or as an AI slop detector API for identifying low-quality, machine-generated content.

The system is trained to be able to handle LLMs from different vendors, such as OpenAI's GPT family of models, Google's Gemini models, Anthropic's Claude models, and the open-release Llama and Mistral models. It is also somewhat robust to small changes and noisy text.

AI detection APIAt a glance
EndpointPOST https://api.sapling.ai/api/v1/aidetect
ReturnsDocument score from 0 to 1, the fraction of the text that reads as AI-generated, per-sentence scores with offsets, per-token probabilities, and an optional HTML heatmap
Input limitsUp to 200,000 characters per request; at least 300 characters recommended
LanguagesEnglish
SDKsPython, JavaScript
PricingFrom $0.005 per 1,000 characters, with volume discounts — see API Pricing

Run Detector

Accuracy and Trade-offs

All AI detection systems have false positives and false negatives. In some cases, small modifications to AI-generated text can cause that text to no longer be flagged as AI-generated. In other cases, human-written (but perhaps rote) text can be misclassified as AI-generated. Please do not interpret this as conclusive belief that your text is AI slop. Depending on the application, false positives or false negatives may be less desirable. Contact us for ways to adjust for your use case.

Quickstart​

  1. Register for a Sapling account and generate a key from your API settings dashboard. See API Access for the full walkthrough.
  2. POST your text to https://api.sapling.ai/api/v1/aidetect with your key, using one of the samples below.
  3. Read score from the response — closer to 1 means more confidence the text is AI-generated. Read ai_fraction to see how much of a document reads as AI-generated (useful when only part of it is), and use sentence_scores or token_probs to show where in the text that comes from.

Sample Code​

curl -X POST https://api.sapling.ai/api/v1/aidetect \
-H "Content-Type: application/json" \
-d '{"key":"<api-key>", "text":"This is sample text."}'
AI Detector UI integration

Sapling's Javascript SDK provides a complete end-to-end UI integration for AI content detection capabilities. Head over to our AI Detect JavaScript Quickstart for more details.

AI Detector POST​

Request Parameters​

https://api.sapling.ai/api/v1/aidetect

HTTP method: POST

The AI Detector API POST endpoint takes JSON parameters documented below:

key: String
32-character API key. Can also be supplied via the Authorization header as a bearer token; if both are provided, the key parameter takes precedence.

text: String
Text to run detection on. The limit is currently 200,000 characters. If latency is high or requests time out, we recommend adapting this script. Please contact us if you need to run the system on longer inputs. We can also provide suggestions on how to chunk your text into smaller pieces and then combine detection results.

Provide exactly one of text and texts.

texts: List[String]
Batch form (see Batch Requests): 1–10 texts to run detection on in one request, each scored independently under the same options.

sent_scores: Boolean
Whether to return sentence scores. Defaults to true. Sentence scores are computed from the same model pass as the document score, so they add almost no latency.

score_string: Boolean
Whether to return string highlighting token-level scores. Defaults to false. This allows you to visualize which portions of the text are likely AI-generated similar to on Sapling's AI detector page.

version: String
There are currently 3 versions of the detector available:

  1. 20240606 (only available upon request)
  2. 20251027 (previous default; available by specifying this version)
  3. 20260820 (current default)

September 17, 2026 default-model upgrade​

The default AI Detector model was upgraded to version 20260820 on September 17, 2026.

The upgrade brings higher precision on modern text and better recall on humanized AI text. In our internal held-out evaluations, human-text false positives fell by 60% on the modern-text set, while detection of humanized AI text improved by 3.9% absolute. We've also refreshed training data from recent ChatGPT, Claude, Gemini, DeepSeek, and Qwen models, and added Kimi coverage.

Requests that omit the version parameter use the new model. To keep using the previous model, include "version": "20251027" in each request.

October 2026 update: ai_fraction and model-based sentence scores​

Two changes to the response, with no change to the request or to score:

  • New ai_fraction field on every response: the share of the text (from 0 to 1) in sentences scoring 0.5 or above. The document score measures how AI-like the text is as a whole, so a document that is half human-written and half AI-generated gets a low score. ai_fraction reports roughly 0.5 for it, so partially AI-written documents are no longer hidden.
  • sentence_scores now come from the detector model itself, on the same scale and with the same 0.5 threshold as score. Each entry also includes start and end character offsets into text for highlighting. Previously, sentence scores were computed by a separate method and could disagree with the overall score. Sentence boundaries may also differ slightly from before.

Requests with default options are also faster, since sentence scoring no longer needs a second model pass.

Response Parameters​

The AI Detector POST endpoint returns JSON of the following format:

{
"score": 0.8016229165451867,
"ai_fraction": 1.0,
"sentence_scores": [
{
"score": 0.8016229165451867,
"sentence": "Here is a sentence.",
"start": 0,
"end": 19
}
],
"text": "Here is a sentence.",
"token_probs": [
0.8062431365251541,
0.8068526238203049,
0.8062431365251541,
0.8080672174692154,
0.8062431365251541
],
"tokens": [
"Here",
" is",
" a",
" sentence",
"."
]
}

A score from 0 to 1 will be returned, with 0 indicating the maximum confidence that the text is human-written, and 1 indicating the maximum confidence that the text is AI-generated.

ai_fraction: The fraction of the text, from 0 to 1, in sentences whose score is 0.5 or above (weighted by characters, ignoring whitespace). Use it alongside score for documents that may be only partly AI-generated: a document with one AI-generated paragraph out of two has a low score but an ai_fraction near 0.5.

If score_string is set to true, a score_string field will be provided. The field contains an HTML string with a heatmap of the portions of the text that are predicted to be AI-generated. If the default score string is not what you desire, you can generate your own using tokens and token_probs.

If sent_scores is true (the default), a field sentence_scores containing scores for each sentence will also be returned. Each entry contains the sentence, its score (which uses the same model, scale, and 0.5 threshold as the overall score), and the start/end character offsets of the sentence in text, so text[start:end] equals sentence (end is exclusive). Offsets count Unicode code points (Python str indexing), not UTF-16 code units or bytes. In JavaScript, slice with Array.from(text) if the input may contain characters outside the Basic Multilingual Plane (e.g. emoji), whose UTF-16 string indices would otherwise misalign. The overall score is computed over the whole text at once, so it is not the plain average of the sentence scores.

tokens: List of tokens from backend tokenizer that can be used to token_probs to visualize the output prediction per token.

token_probs: List of probabilities that each token is AI-generated. This can be used with tokens to visualize the output prediction per token.

Batch Requests​

To score many texts — a queue of submissions, a page of reviews — in one request, send texts (a list of 1–10 strings) instead of text. The same sent_scores, score_string, and version options apply to every item. Exactly one of text and texts must be provided, and the combined length of all items may be up to 200,000 characters (the same cap as a single text).

curl -X POST https://api.sapling.ai/api/v1/aidetect \
-H "Content-Type: application/json" \
-d '{"key":"<api-key>", "texts": ["First submission to check.", "Second submission to check."], "sent_scores": false}'

The batch response is {"results": [...]} with one entry per input, in input order; each entry has exactly the single-response shape above (sentence_scores, tokens, and token_probs are relative to that item's own text):

{
"results": [
{"score": 0.93, "text": "First submission to check.", "tokens": ["..."], "token_probs": [0.9]},
{"score": 0.04, "text": "Second submission to check.", "tokens": ["..."], "token_probs": [0.1]}
]
}

Each item is scored independently and billed individually (a batch of N costs the same as N single requests, and identical items are billed once). If any item fails, the whole request returns an error and nothing is billed; items that did resolve are cached, so a retry only re-scores the rest. The batch form requires an API key.

Interpreting the score​

The detector is a probability estimate, not a verdict. How you turn score into a decision depends on which kind of mistake is more costly in your application:

  • A high threshold (for example, only act above 0.9) minimizes false accusations, at the cost of letting more AI-generated text through. This is usually the right default when a flag has consequences for a person — moderation, admissions, hiring.
  • A lower threshold surfaces more candidates for human review. This suits content pipelines where a flag just means "route to an editor".
  • ai_fraction and sentence scores are useful for review workflows: a document that is mostly human-written with a few AI-generated paragraphs has a low overall score, but its ai_fraction and the flagged entries in sentence_scores show how much of it was AI-generated, and where. Consider flagging on ai_fraction (for example, above 0.3) when partially AI-written submissions matter to you.

Contact us if you need help calibrating thresholds against your own labelled data.

Limits, quotas, and pricing​

  • Input length: up to 200,000 characters per request. At least 300 characters is recommended — it is possible to classify shorter text, but there is much less signal to work with.
  • Free quota: new API keys get 50,000 characters/day and 250,000 characters/month. See Usage and Pricing.
  • Rate limits: 120,000 characters every 2 minutes for this endpoint, plus the daily and monthly quotas above. See Rate Limits; status code 429 indicates a rate limit error.
  • Caching: AI detection results are cached on the entire text for 3 days, so re-sending an identical document does not incur additional usage. See API Pricing.
  • Billing: usage-based, per character, starting at $0.005 per 1,000 characters with volume discounts. See API Pricing.

Tips​

Checking Files (PDF/DOCX)​

Sometimes you may wish to send the API PDFs or DOCX files.

To do this, refer to the Files documentation to see how you can extract text from files before passing the text to the API. These endpoints are currently provided free-of-charge; however, if you plan to use them for high-volumes of text, contact us and ensure you're using one of the other endpoints or we may limit usage to reduce server load.

Other ways to use Sapling's AI detection​

Beyond calling the HTTP API directly, the same detector is available through:

  • JavaScript SDK — a drop-in UI that adds an AI detection button to any textarea or contenteditable in your web app.
  • Python SDK — the sapling-py package wraps the endpoint as client.aidetect(...).
  • Sapling MCP Server — exposes AI detection as a tool to Claude and other MCP-compatible assistants.
  • Sapling's AI content detector — the hosted web app, useful for spot checks and for comparing against your own integration.

If you are still evaluating vendors, Sapling maintains a comparison of the top AI detection APIs.

Frequently asked questions​

What is an AI detection API?​

An AI detection API is an HTTP endpoint that takes text as input and returns a score estimating how likely that text was generated by a large language model rather than written by a person. Sapling's AI detector API returns a document-level score from 0 to 1, plus optional per-sentence and per-token scores so you can see which passages drive the result.

Which AI models can the detector identify?​

The detector is trained across LLMs from different vendors, including OpenAI's GPT family, Google's Gemini models, Anthropic's Claude models, and the open-release Llama and Mistral models. It does not name which model produced the text; it returns the probability that the text is AI-generated.

How much text does the AI detector API need?​

At least 300 characters is recommended. Shorter inputs carry much less signal, so scores on them are less reliable. A single request accepts up to 200,000 characters.

How accurate is AI detection?​

Every AI detection system has false positives and false negatives. Small edits to AI-generated text can keep it from being flagged, and human-written but rote text is sometimes misclassified as AI-generated. Treat the score as a signal to review rather than as proof, and contact us to tune thresholds for your use case.

Does the AI detector API support languages other than English?​

The AI detector is currently trained for English only. You can use Sapling's language detection endpoint to route non-English text elsewhere in your pipeline.

Can I run AI detection on PDF or DOCX files?​

Yes. Extract the text with Sapling's PDF-to-text or DOCX-to-text endpoints, then pass the extracted text to the AI detection endpoint.

Is there a free tier for the AI detection API?​

Yes. A new API key comes with a free quota of 50,000 characters per day and 250,000 characters per month, which is enough to evaluate the detector before subscribing. Paid usage is billed per character with volume discounts — see API Pricing.