logo

Keep Both Versions: Translate PDF Annotations for Language Learners

·by WordByWord Team
10 min read
Keep Both Versions: Translate PDF Annotations for Language Learners

The simplest way to translate PDF annotations and keep them attached to your highlights is to use an annotation-aware plugin that saves the translation directly into the comment or note, or into a paired sidecar file. For scanned pages, run OCR first since the text has to exist before any tool can translate it. For a full translated copy of a document, a whole-file translator works, though cloud tools may not edit the original file in place.


TL;DR:

  • Store translations in annotation comments to keep them attached to highlights; sidecar files can preserve layout, but must stay alongside the PDF.
  • Scanned PDFs need OCR before translation; Google Drive works best with files no larger than 2 MB and text at least 10 pixels tall.
  • Whole file translation suits readers needing a complete copy, but annotations may not transfer; keep the original and check page count, file size, and layout support.
  • Cloud assistants can translate batches of annotation text, but they may not preserve PDF highlight geometry; avoid uploading sensitive files without checking data handling.

WordByWord
Build Vocabulary From PDFs
Translate unknown words in PDFs, save them to collections, and review them with spaced repetition as you read.
Explore WordByWord

Table of Contents

How annotation-aware plugins keep translations tied to your highlights

Annotation-aware PDF tools work on a simple loop: you select text, trigger a translation, and the result gets written into the annotation’s comment field or into an attached note rather than replacing the original passage. Reference implementations like zotero-pdf-translate show this pattern clearly: translating a selection can add the result to the annotation comment or body, or create a separate note that holds both the source and the translation side by side.

Before you rely on one of these tools, check a few settings:

  • Whether automatic annotation translation is on, so new highlights get translated without a manual step.
  • Whether there is an option to add the translation directly to the note rather than overwrite the highlight.
  • Whether a delimiter separates source and target text, which matters if you plan to retranslate later.
  • Whether the tool supports updating or retranslating an existing annotation without duplicating it.

Persistence matters here too. Saving the translation inline, inside the comment body, usually survives exporting comments to another reader. A paired sidecar file (sometimes called an overlay or a .translations.md file) can preserve page coordinates and formatting better, but it only works if you keep that file alongside the PDF.

Pro Tip: Before installing a PDF extension, check what permissions it requests. A tool that only needs access to the active PDF tab is a safer bet than one asking for broad browsing history access.

OCR, translate, reinsert: a workflow for scanned PDFs

Scanned or image-only PDFs have no underlying text layer, so a translation tool has nothing to read until OCR converts the pixels into characters. Google Drive’s OCR handles this conversion but comes with real constraints: it works best on files 2 MB or smaller, text needs to be at least 10 pixels tall, pages need correct orientation, and the image needs to be sharp. Tables, columns, and multi-part layouts often don’t survive the conversion cleanly.

A practical sequence looks like this:

  1. Run OCR on the scanned PDF, keeping file size and resolution within the limits Google Drive recommends.
  2. Extract the resulting text and scan it for obvious OCR errors, especially around tables or footnotes.
  3. Translate the specific passage or annotation you need, rather than the whole page, to make proofreading faster.
  4. Paste the translation into an annotation comment or a note attached to the original highlight.
  5. Compare the translated text against the original image to catch any OCR misreads that slipped into the translation.

Google Drive’s OCR works most reliably on clean, well-lit scans under its stated size and resolution limits, and complex layouts like tables or multi-column text are the most common failure point.

Once you’ve translated the passage, decide where it should live. An annotation comment keeps it attached to the highlight inside the same file. A sidecar bilingual file keeps source and translation separate but searchable together. A bilingual exported PDF, with both languages visible, works well if you’ll be sharing the document with someone else who doesn’t need the annotation workflow at all.

Three ways to store a PDF translation

When to translate the whole file instead of individual annotations

Document-level translators take a different approach: you upload the entire PDF and get back a fully translated copy. Google Translate’s document feature supports this upload-translate-download flow for PDFs and other formats, though it notes limitations on very large or very long documents, so check page count and file size before you start.

Whole-file translation makes sense when you need a complete readable copy, say for sharing with someone who reads only the target language. Per-annotation translation makes more sense when you’re studying from the original and just need specific passages or notes translated for yourself.

If you translate the entire file, your original annotations usually don’t carry over automatically. A few ways to keep them useful:

  • Export and keep the original annotated PDF separately, so your highlights and notes stay intact.
  • Save the full translation as a linked bilingual copy you can reference alongside the original.
  • Keep a sidecar translation file that maps specific annotations to their translated text, even after the full document translation is done.

Before choosing a whole-file tool, confirm its file-size limit, its page-count limit, and whether it preserves layout elements like tables and columns, since these are the most common points of failure in document translation.

Using an LLM to extract and translate annotation text

Document assistants like ChatGPT can accept PDF uploads and extract text for translation, which works well for bulk-translating a batch of annotation text at once. OpenAI’s documentation notes that PDFs are a supported file type, but visual retrieval, meaning how well the tool reads layout and embedded images, is plan-dependent, with richer visual handling reserved for enterprise-tier features.

A workable pattern:

  • Export the annotations or the annotated passages as plain text.
  • Send that text to the assistant with a clear translation request.
  • Paste the translated result back into the annotation comment or note in your PDF reader.

Pro Tip: If the PDF contains sensitive personal or financial information, keep the translation step local or confirm the tool’s data handling policy before uploading, rather than defaulting to a cloud assistant.

LLMs are strong at translating selected chunks of annotation text in bulk. They’re a poor fit for editing the PDF’s actual geometry, repositioning highlights, or preserving bounding boxes, since standard plans generally don’t edit the original file’s annotation layer directly.

How WordByWord helps you translate and retain annotated text in PDFs

Reading a PDF in another language usually means stopping on a word, losing your place, and either skipping it or digging through a separate translator tab. Our tools enable selecting a word or phrase inside the PDF and getting an inline translation in one click, across many languages, without leaving the page.

A Spotlight mode highlights unfamiliar vocabulary directly on the page based on your familiarity with each word, so you can see at a glance what’s worth your attention. A comprehension indicator can show what share of the words on the page you know before you commit to reading it closely.

A suggested workflow is to look up a phrase inline as you read, save it to a collection, and if you want that translation embedded in the file itself, copy it into the PDF’s native comment feature as a short note. Saving translated phrases into a collection, rather than translating once and moving on, can help turn a one-time lookup into vocabulary you retain, as saved words are reviewed through spaced repetition.

What actually works best for most readers

Most guides treat translation tools as interchangeable, but the right choice depends entirely on whether you need the translation to stay attached to your reading or just need a finished copy. Our recommendation: default to an annotation-aware local workflow whenever the tool supports it, since it keeps your highlights intact and your translations editable in place.

For scanned documents, run OCR first and verify the output before trusting any translation built on top of it. Reserve cloud LLMs and full-file translators for documents that aren’t sensitive, since you’re sending complete content to a third party. In every case, save the translation somewhere retrievable, a comment or a paired file, rather than overwriting the source text you might need to check again later.

— WordByWord Team

Try WordByWord for inline PDF translation and vocabulary that sticks

If you’re translating PDF annotations because you’re learning a language, not just reading a document once, we built our tools for exactly that pattern. Instead of translating a passage and forgetting it, you look up a word inline, see it highlighted by how well you know it, and save it to a collection that resurfaces it through spaced repetition until it sticks.

WordByWord

Getting started takes three steps:

  • Install our browser extension or open the web app and open your PDF.
  • Translate words or phrases inline as you read, in any of 50+ languages.
  • Save new vocabulary to a collection so it carries over into review sessions instead of disappearing.

Our Free Forever plan has no time limit, so you can test the inline translation and collection workflow on your own PDFs right now.

FAQ

How do I translate a PDF automatically?

Upload the PDF to a document translator such as Google Translate’s document feature, which translates the full file and gives you a downloaded copy. For translations tied to specific highlights instead of the whole file, an annotation-aware plugin that saves translations into comments is a better fit.

Why can’t I translate a scanned PDF directly?

A scanned PDF is just an image, so there’s no text layer for a translation tool to read until OCR converts it to actual characters. Google Drive’s OCR handles this conversion but works best on sharp images under its recommended size and resolution limits, and complex layouts like tables often don’t convert cleanly.

Can ChatGPT translate an entire PDF?

ChatGPT can accept PDF uploads and extract text for translation, which works well for bulk-translating annotation text, according to OpenAI’s documentation. Visual retrieval and how well layout and embedded images are preserved depend on the plan, with richer handling reserved for enterprise features.

Does Adobe Acrobat have a built-in translate feature?

Acrobat doesn’t translate text automatically, but it does let you add and edit comments attached to selected text, which is a practical place to store a translation you’ve done elsewhere, per Adobe’s comment documentation. Many readers pair a separate translation step with Acrobat’s comment tools to keep the result attached to the original highlight.

How do I keep translated annotations working across different PDF readers?

Saving the translation into the annotation’s comment field, rather than a separate app-specific layer, usually survives exporting between readers. For cases where layout and coordinates matter, a sidecar file format alongside the PDF offers more reliable persistence across devices and exports.

Sources

Request a feature

What should we add or fix? We read every message and send a little gift for the idea.

0 / 4000