Document File Size Limits
We impose size limits per file type. Limits vary by format and API plan. For DOCX and PPTX, our .NET, PHP and NodeJS client libraries offer functionality to minify files by temporarily extracting large media formats before sending them to the DeepL API. Stripped media are reinserted after document translation is completed. This allows users to translate files that might hit the size limit.Billing minimums
Every submitted document of type.pptx, .docx, .doc, .xlsx, or .pdf is billed a minimum of 50,000 characters on DeepL API plans, no matter how many characters the document contains. Formats currently in beta are not billed and are not subject to this minimum.
One source/target language pair per upload
Thesource_lang and target_lang values on the request apply to the entire uploaded file. For most formats, keep each upload to a single source language for consistent results — behavior on content that isn’t in the selected source language is not guaranteed.
XLIFF is the exception: <source> elements are translated as independent segments, so XLIFF files containing segments in different languages are handled segment-by-segment. Note: the per-<file> source-language attribute inside an XLIFF is ignored — DeepL uses the request’s source_lang value for every segment.
All content in the uploaded file counts toward billed characters, including content that was not actually translated.
Same-language source and target are rejected
Requests where the source and target languages are equal — including regional variants of the same language, e.g.EN → EN-US or EN-US → EN-GB — are rejected with HTTP 400 (Source and target language are equal.). To adapt regional spelling, post-process the translated output yourself.
Format-specific gotchas
Each supported format has behaviors and constraints worth knowing before you upload. The most common surprises: XML- Only text between element tags is translated. Attribute values and processing instructions are left alone.
- XML
<!-- comments -->may be picked up as translatable text and their delimiters escaped in the output — strip comments before uploading if downstream tooling relies on them. CDATAsection content is translated — the markers are stripped and the inner text is sent to the engine. Move code, regex, and other non-translatable content out ofCDATAblocks first.translate="no"on any XML element excludes its content from translation.- Files using ITS 2.0 external rules (
its:rules/@xlink:href) currently return HTTP 500 — inline the rules or remove the reference. - Malformed XML currently returns HTTP 500 instead of a 4xx. Validate with a standard XML parser before uploading.
- Only
<source>content is translated; results are written to the corresponding<target>. Existing<target>values are overwritten — remove or split out already-translated units if you need to preserve them. - The per-
<file>source-language/target-languageattributes are ignored — DeepL uses the API’ssource_langandtarget_langfor the entire upload. translate="no"on<trans-unit>(1.2) or<unit>(2.0) is fully honored.- The
stateattribute is preserved as-is; if you usestate="needs-translation", update it yourself after translation. - An unsupported
trgLangvalue on the root<xliff>element (e.g.arb-MOD) is rejected as “Invalid target language” even if the APItarget_langis valid — remove or fix the attribute.
- Only DITA topic files (
.dita) are supported..ditamapuploads are rejected. translate="no"is fully honored on inline and block-level elements.- Content references (
conref,conkeyref) are not resolved — translate each referenced source topic independently. - Newlines inside
<codeblock>,<pre>,<msgblock>, and<screen>may be collapsed to single spaces. Move code samples out of the file before translation if line breaks matter. - Topics with dense inline-element markup (combinations of
<filepath>,<cmdname>,<option>, nested<indexterm>) may fail with an internal error — simplify or split the topic.
- Only string values are translated; keys, numbers, booleans, and
nullare not touched. Nested objects and arrays are traversed at every depth. - Files must be strict, parseable JSON — no trailing commas, no comments. JSONC-style extensions are not supported.
- Upload limit is 1 MB regardless of plan. Large metadata payloads (e.g., DataCite, Zenodo, Backstage catalog dumps) may need to be split, or translated string-by-string via the text-translation API.
- Embedded HTML or Markdown inside string values (common in Contentful Rich Text and similar CMS payloads) is handled — DeepL translates the natural-language text and attempts to preserve the embedded markup. Review the output for complex rich-text content.
- Brace-delimited placeholders such as
{stars}or{{userName}}are kept verbatim by default, so template variables survive translation. Multi-word groups such as{see note}are treated as translatable text. To translate placeholders along with the text, setinput_conversion_options=version:1,json-placeholders:translatewhen uploading. - ICU MessageFormat skeletons such as
{count, plural, one {# item} other {# items}}are always protected: the variable name, keywords, and categories stay intact while the branch text is translated. Branch text is translated in isolation, so review plural forms in the output. Languages whose plural categories are absent from the source (for example Polishfew/manyfrom an English message) fall back toother. Expand categories in your i18n tooling if exact pluralization matters. - To protect specific values from translation, encode them as non-strings (numbers/booleans/null) or pre-process the file to strip them.
- InDesign embeds font references, not the fonts themselves. If the target language uses characters not in the original font (e.g., Japanese in a Latin-only font), the output may show boxes or substituted glyphs — open the translated file in InDesign and swap fonts before distributing.
- Translated text is often longer than the source. Expect overset text (indicated by a red ”+” in InDesign) in fixed-size frames; review and resize after translation.
- IDML does not use
translate="no". Protect content via InDesign character styles or by removing the affected text frames before exporting.
- The file must be genuine Adobe FrameMaker MIF.
.miffiles from other tools (e.g. Quartus memory-init files, MathML wrapped as MIF) currently return HTTP 500 rather than being rejected cleanly. Rename or convert them before uploading. - MIF 8.00 and later are supported. Older MIF variants are best-effort — open and re-save from a recent FrameMaker version if the upload fails.
- Save as UTF-8 from FrameMaker before uploading; the file should begin with a
<MIFFile ...>header. - Some valid FrameMaker 10 MIF files may fail with HTTP 500 — try re-saving from a newer FrameMaker version.
- MIF does not use
translate="no". Protect content via FrameMaker conditional text or character formatting.
- Cell text is translated across all sheets; rich text formatting, merged cells, and multi-sheet layouts are preserved. Numbers and dates are left unchanged.
- Macros are preserved byte-for-byte and never translated. The VBA project is carried through untouched, so user-facing strings inside macro code remain in the source language.
- Formulas are preserved verbatim and cached values are kept as-is.
- Lookup formulas (
VLOOKUP,MATCH,INDEX) whose lookup key is a text literal typed into the formula keep that literal in the source language while the table they search is translated, so the lookup can stop matching and return#N/Aafter recalculation. Key lookups off a cell reference (e.g.=VLOOKUP(D1,...)) instead. - Files with no translatable text are rejected with
No translatable text can be extracted from the document.This does not mean the file is malformed.
.zip) (currently in beta)
- The
.zipmust be a valid SCORM package (SCORM 1.2 or SCORM 2004) withimsmanifest.xmlat the root of the archive. AICC, xAPI (Tin Can), and cmi5 packages are not supported and are rejected. - Course titles in
imsmanifest.xmland lesson content in the package’s HTML files are translated. Manifest identifiers, file paths, and HTML markup are preserved exactly. - Audio and video files (MP3, MP4, WAV, images) are preserved byte-for-byte. DeepL does not transcribe or translate media content.
- Zip the contents, not the folder.
imsmanifest.xmlmust sit at the root of the archive, exactly one. An archive that starts with a folder (e.g.course/imsmanifest.xml) is rejected. - The manifest must be under 4 MB. A larger
imsmanifest.xmlis rejected. - Non-SCORM
.zipuploads are rejected with400 Invalid or missing file extension.
- Cue text is translated; cue timings, identifiers, settings lines, and
WEBVTTheaders are preserved so subtitles stay in sync with the original media. - Inline styling tags (e.g.
<c>,<v Speaker>, karaoke timestamps) are preserved. Review their placement, since translated text length differs from the source. - Files with no translatable text are rejected.
- String values are translated at every level of nesting; keys, numbers, booleans,
nullvalues, comments, and indentation structure are preserved. - Placeholders and escape sequences (
\n,\t, etc.) are preserved. - Files with no translatable text are rejected.
- Malformed YAML (inconsistent indentation, unquoted special characters) returns an error. Validate before uploading.
.properties (currently in beta)
- Property values are translated; keys and separators are preserved. Comment lines (
#or!) are not translated. - Placeholders (
{0},%s,%d) are normally preserved. Verify them after translation before shipping. - Escaped characters (
\n,\t,\:) are preserved. - Files with no translatable text are returned unchanged.
.strings (currently in beta)
- String values are translated; keys and the
"key" = "value";structure are preserved. - Format specifiers (
%@,%d, positional specifiers like%1$@) are normally preserved. Verify them after translation before shipping. - Escaped characters (
\n,\",\\) are preserved. - Files with no translatable text are returned unchanged.
.md / .markdown) (currently in beta)
- Paragraphs, headings, list items, blockquotes, table cell content, and image alt text are translated. URLs and link targets are not, only the link label text is.
- YAML front matter is not translated. Known issue: the
---delimiters that open and close the front matter block are currently dropped from the translated file, which can fuse the metadata into the first paragraph. Re-add the---lines after download, or move the front matter out of the file before uploading. - HTML blocks embedded in Markdown are processed via an HTML sub-filter. Results may vary for complex inline HTML.
- Markdown formatting (bold, italic, headings, lists) is preserved.
- Malformed Markdown is accepted and translated as-is (not validated for well-formedness).
- String values in
<data>elements are translated; element names, IDs, and XML structure are preserved. Comments in<comment>elements are not translated. - Placeholders (
{0},%s) are normally preserved. Verify them after translation before shipping. - Files with no translatable text are returned unchanged.
- Body text, headings, table content, footnotes, endnotes, annotations, and document metadata are translated.
- Accept or reject all tracked changes before uploading. Deleted text that has not been formally removed may still be extracted and translated.
- Consecutive tabs may be dropped during translation. Review content that relies on tab-based alignment.
- If the target language requires characters outside the document’s fonts, they may not render correctly. Check fonts after translation.
- Body text in paragraphs, headings, list items, table cells, footnotes, and endnotes is translated. RTF control words, groups, and formatting tokens (
\b,\i,\par,\fonttbl, etc.) are preserved untouched. - Non-ASCII characters are written as
\uXXXX?escape sequences. Old RTF readers that ignore\uescapes will show only the ASCII fallback character. Open the result in a modern reader (Word 2007+, LibreOffice). - Embedded objects (images, OLE objects, embedded fonts) are preserved as binary blocks and not modified.
- Fields (
\field) are preserved, but field instruction text is not translated. Only the displayed text result is. - Hyperlink targets are not translated; only the link’s display text is.
Polling and translation time
Translation time depends on document size and server load: small documents typically finish in seconds, larger ones in 1-2 minutes once translation has started. Poll the status endpoint at regular intervals or with exponential backoff. Treat theseconds_remaining field as a rough estimate only; it can be unreliable and occasionally returns implausible values (e.g. 2^27).
Using glossaries with documents
You can apply a glossary to a document translation with theglossary_id parameter (or up to 5 glossaries with glossary_ids). This requires the source_lang parameter to be set, and the glossary’s language pair has to match the language pair of the request.
Requesting a quality evaluation
A quality evaluation is a report on the translation DeepL just produced: for each segment of the document it lists the issues it detected, such as mistranslations, omissions, and fluency or style problems, with a severity and the character range each issue applies to. Use it to target human review at the segments that need it instead of reviewing the whole document. For a product overview, see About quality evaluation for file translations. To request a quality evaluation, setenable_quality_evaluation to true when uploading a document.
When enable_quality_evaluation=true, three additional checks run before the translation request is accepted:
- The language pair must be in the supported language pairs. Quality evaluation supports fewer language pairs than document translation.
- The file must be
docx,pptx,pdf,idml,xml,dita,mif, or XLIFF 1.2/2.0/2.1. Quality evaluation supports fewer file types than document translation. - Your account must have enough quality evaluation characters. Quality evaluation characters are metered separately from translation characters.
quality_evaluation_job_id alongside the usual document fields. It is the unique ID of the evaluation, and the value you poll with to retrieve the report:
"status": "processing"; wait the number of seconds in the Retry-After header and poll again. Expect this for a while after the translated document is ready, because the evaluation only starts once there is a translation to evaluate.
A finished evaluation returns HTTP 200 whether it succeeded or failed, so branch on status rather than the status code: done carries the report, and error means the evaluation could not be produced, for example because the language pair isn’t supported.
done: reports are kept for 24 hours after the evaluation finishes, and polling after that returns HTTP 404. For the severity and issue-type vocabulary, how to resolve source_spans and target_spans, and the applied glossary term pairs a report returns per segment, see the quality evaluation reference.
Document format conversions
By default, the translated document comes back in the same format as the input. Two conversions differ:- Translating a
.docfile returns a.docxfile. - With the
output_formatparameter on upload, you can translate a PDF and receive an editable Microsoft Word document (output_format=docx). No other input formats support alternative output formats.
Error 429: Too Many Requests
This error may occur when:- You send concurrent document translation requests that exceed your account quota.
- You have too many un-retrieved or non-downloaded translated documents.
- Polling document translation status, taking into account its frequency in order to prevent excessive load
- Quicker time to document retrieval
- Retries with exponential backoff