sample1_2.pdf
Hi,
Thank you for the detailed response. Here are the requested details:
Document Translation API version: v1.1
I am using the Microsoft Translator V3 connector in Power Automate. The operation-location header returned by the service confirms the endpoint being called:
https://...cognitiveservices.azure.com/translator/text/batch/v1.1/batches/{id}}
Workflow: Asynchronous (batch). The flow calls StartDocumentTranslation, receives an operationID, then polls GetDocumentsStatus until the document status is Succeeded or Failed.
Source and target languages: Source is set to Auto-detect, since our documents often contain more than one language. Target varies per request (en, de, fr, it and 12 others), selected by the user from a SharePoint metadata column. For the purpose of test I will provide all documents in italian translated to english.
PDF content: Both. The document contains selectable text and embedded images. I verified this with Ctrl+F in the PDF: the text inside the title block tables is searchable, while the company logo is not confirming the logo is an image, not text.
Sample files: You can see the documents attached to the reply.
Additional question about the 2026-03-01 API version
You mentioned that the 2026-03-01 Document Translation API supports PDF translation using Azure Document Intelligence with the goal of preserving PDF layout and structure. Two questions:
- Would this version be expected to resolve the image/logo re-rendering I am seeing, or does it only address text layout reflow?
Is this version accessible through the Power Automate Microsoft Translator V3 connector, or does it require calling the REST API directly?
Summary of the observed behavior
I will share two sanitized sample sets, because they behave differently and the comparison may be useful:
Sample 1: native PDF exported from DWG (CAD drawing)
Never scanned
Fonts changed throughout the document, including in areas that were not translated
The company logo, which is an embedded image and not searchable text, was re-rendered with a substitute font.
Sample 2: native PDF exported from WORD
Also shows font substitution and text shifting
However, the embedded images are fully preserved and untouched
Sample 3: WORD version of Sample 2
This version is to show that when it is not PDF, the translation is completed perfectly.
In Sample 2 the images behave exactly as the FAQ describes: digital portions translated (with a different font), images left alone. In Sample 1 the logo is clearly re-rendered despite being an image. This is why I suspect the processing path differs depending on PDF structure, possibly with OCR being applied to the DWG-derived file.
LAST UPDATE: While I was trying to prepare a sample of the PDF derived from a .dwg file, I translated it multiple times and the problem of the logo seemed resolved, but i do not understand why. I will still provide that file as a sample no1, and I appreciate if you can explain the main reason why it got fixed, so I can use the same approach for translating documents like that.
That being said, I still lose the font, and the information like headers, titles (the ones that are written with bold/italic) in native PDFs. I would appreciate it, If you can provide me an approach I can use for native PDFs that allow me to provide an optimal output.
Thank you for your help.