wxrks TMill is wxrks' AI-powered service for cleaning up large, aging Translation Memories (TMs) — it uses semantic analysis to find and fix quality problems that simple syntax or keyword checks miss.
💡 Who is this for? This guide is for Account Admins who need to request a semantic clean-up of a large or degraded Translation Memory within the wxrks platform.
Why Translation Memories need cleaning up
A TM gets more valuable the longer it's used — but it also accumulates errors over time. On large, long-running TMs, this can add up to a wide range of quality problems, including but not limited to:
Major-severity issues
A mismatch between the declared locale and the actual language in a Translation Unit Variant (TUV) — the source or target entry within a saved segment pair.
Example: a TUV labeled PT-BR actually contains Korean text.Incorrect translations in TUVs that significantly deviate from the source meaning.
Minor-severity issues
Tag mismatches between source and target TUVs.
Terminology mismatches between TUVs and the associated glossary.
Spelling errors.
Grammar mistakes.
Outdated tone (e.g., formal vs. informal).
Culturally inappropriate language (severity can vary by context).
Left unaddressed, these problems quietly reduce how much leverage you actually get from the TM, and can propagate bad translations into new projects that reuse the same matches.
Why ordinary validation isn't enough
Most of the issues above can't be caught by syntactic validation alone — they require semantic analysis, understanding what a segment actually means, not just whether it's well-formed. That kind of analysis has traditionally been too costly to run across an entire TM. As a result, teams have typically dealt with degraded TMs in one of two ways: applying a TM-wide penalty (which reduces match leverage and increases translation cost), or tolerating the errors and letting them keep propagating into new translations.
How wxrks TMill works
TMill is wxrks' semantic TM clean-up service, powered by the wxrks ML tech stack. It combines models including GPT-3.5 and GPT-4 with proprietary NLP tooling to clean TMs of any size, locale, or condition.
What you'll need: a TMX export of your TM (see How to export the Translation Memory).
Minimum processing unit per request: 1,000,000 TUVs.
Note: a self-service version of TMill is in development inside wxrks Labs, so you'll eventually be able to run smaller clean-ups directly from the app. Until it ships, every TMill clean-up runs as a guided engagement with the wxrks team — see "Request a TM clean-up" below.
Stage 1: Asset preparation
After you send your TMX file(s), wxrks runs integrity checks to confirm they're ready for processing — locale consistency, the absence of a high number of empty TUVs, and sound file structure. Any critical issues are flagged before moving forward, so problems in the file itself don't get mistaken for problems in the translations.
Stage 2: Initial analysis
Once your TMX passes preparation, the TMill engine scans every TUV for the error types listed above and produces a report showing the total number of TUVs analyzed and how many fall into each error category.
Note: if a TUV matches more than one error category, it's grouped under whichever category has the highest severity.
Stage 3: Joint strategy definition
Based on the analysis, wxrks shares clean-up recommendations and works with you to define the final approach — some errors may be flagged as non-recoverable and removed outright, while others can be corrected. The plan is shaped around your specific use case, not a one-size-fits-all fix.
Stage 4: Clean-up execution
The engine reprocesses every flagged segment by error type and applies best-attempt fixes. Some fixes (spelling, grammar) tend to be highly accurate; others (like tag corrections) depend more on the file's structure and locale specifics, so results can vary segment to segment.
Stage 5: Packaging & deliverables
The cleaned TM is reassembled according to the strategy agreed on in Stage 3. Depending on what you need, it can be delivered as a single combined master file, or segmented by error type, author, time frame, or other metadata — giving you a foundation for a more strategic approach to TM management going forward.
Request a TM clean-up
Want to see how much value a semantic clean-up could unlock in your legacy TMs? Schedule a conversation with the wxrks team — they'll review your TMX file(s) and walk you through what a clean-up would look like for your specific TM.
