How to create a TMX translation memory from bilingual PDFs

You have a document and its translation, both as PDFs, and you want the sentence pairs in your CAT tool's translation memory. This guide covers what alignment is, how to prepare your files, how to run it with PDF Aligner, how to review the result, and how to import the TMX.

What alignment means

Alignment matches each sentence in a source document with its counterpart in the translation. The result is a list of translation units (source sentence, target sentence) saved as TMX, the standard exchange format for translation memories. Most CAT tools, including Trados Studio, memoQ, OmegaT and Phrase, can import it, so the translations you already paid for become fuzzy and exact matches in your next project.

Text-based aligners struggle when a PDF has no usable text layer, or when columns and tables scramble the reading order. PDF Aligner reads every page with vision AI instead, so scanned and complex layouts work too.

Before you start

  • Two PDFs: the original and its translation, each covering the same content.
  • Matching pages: pages are aligned position by position (page 1 with page 1, and so on), so the best results come from PDFs with matching pages. If the two files paginate differently, trim both to the same section first, or use the page-range option.
  • Start small: run a first test on two or three pages. It costs little and shows you how the model handles your layout before you commit a whole document.
  • Privacy: uploaded PDFs are deleted right after processing. Only the aligned sentence pairs stay in your account so you can review and export them.

Step by step

  1. Create a free account. The free plan includes 5 pages every month, no credit card. Sign up here.
  2. Start a new alignment job. Upload the source PDF and the target PDF, then choose the language pair. Any pair works, not only Japanese and English.
  3. Pick a model. Lower-cost models suit clean, straightforward documents. Stronger models cost more credits per page and handle dense text, difficult scans and non-Latin scripts better. See the model comparison.
  4. Optional settings. Set a page range for your test run. On paid plans, batch processing halves the credit cost in exchange for an asynchronous result.
  5. Run the job. Behind the scenes it works in stages: page OCR, sentence alignment, then clean-up (merging broken fragments, splitting multi-sentence pairs and removing duplicates).
  6. Review the pairs. Each pair gets a quality score. Filter to low-quality (under 50) or medium-quality (50 to 80) pairs and check those first. Use Split, Merge and Delete to fix them, and Delete empty pairs to remove rows that have no source or no target text.
  7. Export. Download TMX for your CAT tool, or CSV or XML if you need another format.

Importing the TMX

Use your CAT tool's translation memory import function and select the file. Two practical points: make sure the language codes match your project (the TMX uses codes such as ja and en; if your tool expects regional variants such as en-US, map them during import), and import into a new, empty memory first so you can check the result before mixing it with existing data. OmegaT reads TMX files placed in a project's tm folder.

Tips for cleaner results

  • Use the same version of the document in both languages. A revised source against an older translation leaves gaps.
  • Remove pages that exist in only one file, such as a cover page or a table of contents, before aligning.
  • For scans or unusual scripts, choose a stronger model rather than re-running a cheap one.
  • Review the low-quality pairs first instead of reading everything. That is where most corrections are.
  • Expect to review. AI alignment is fast but not perfect. In our published sample, 46 of 47 pairs agree with the Ministry of Justice's own TMX of the same text (the remaining difference is a date line the reference merges into the title), and one pair was fixed by hand with the editor.

Frequently asked questions

Can it align scanned PDFs?

Yes. Every page is read with vision OCR, so a missing text layer is not a problem. Poor scans align best with a stronger model.

Which language pairs are supported?

Any pair. You choose the source and target language for each job.

What does it cost?

The free plan covers 5 pages every month. Paid plans add credits, and each page costs a number of credits that depends on the model you choose. See pricing for the current details.

See the output before you sign up

The sample page shows a real Japanese-English result you can download as TMX, no account needed.

View the sample Sign up free