如何通过Google文档翻译API移除文档中的‘Machine Translated by Google’标识?
Google文档翻译API归因标识问题解决方案
一、customizedAttribution参数报错原因
Google文档翻译API对自定义归因有强制规则:自定义内容必须明确提及Google翻译的归属,不能使用无意义内容(如"Test")。合规的自定义归因示例包括「翻译由Google提供支持」「Powered by Google Translate」等,必须保留Google的品牌关联,否则会触发INVALID_ARGUMENT错误。
二、能否移除文档中的归因标识?
付费版API同样要求保留归因,这是服务条款的强制要求。除非你与Google签订了特殊企业定制协议,否则无法通过API直接移除文档内的归因标识——网站声明不能替代文档本身的归因要求。
三、Node.js中优化PDF归因处理的方案
如果需要调整标识(而非完全移除),可以使用pdf-lib库实现更精细的PDF操作,替代简单的白色矩形覆盖:
步骤1:安装依赖
npm install pdf-lib
步骤2:修改现有代码,添加PDF处理逻辑
const crypto = require('crypto') const { TranslationServiceClient } = require('@google-cloud/translate').v3 const { Storage } = require('@google-cloud/storage'); const fetch = require('node-fetch'); const fs = require('fs') const { PDFDocument } = require('pdf-lib'); // 新增依赖 const translationClient = new TranslationServiceClient(); const parent = translationClient.locationPath('***', 'global'); const storage = new Storage(); async function translateFile({ uri, from, to }) { process.env.GOOGLE_APPLICATION_CREDENTIALS = './***.json' const pdfResponse = await fetch(uri); if (!pdfResponse.ok) { throw new Error(`Failed to download the PDF: ${pdfResponse.statusText}`); } const buffer = await pdfResponse.buffer(); const bucket = storage.bucket('***'); const file = bucket.file(`${crypto.randomUUID()}.pdf`) await file.save(buffer) const inputUri = `gs://${bucket.name}/${file.name}` const documentInputConfig = { gcsSource: { inputUri } }; // 使用合规的自定义归因 const request = { parent, documentInputConfig: documentInputConfig, sourceLanguageCode: from, targetLanguageCode: to, customizedAttribution: '翻译由Google提供支持' // 符合规则的自定义内容 }; const [response] = await translationClient.translateDocument(request); await file.delete() // 处理PDF,替换或调整归因标识 const pdfBytes = response.documentTranslation.byteStreamOutputs[0]; const pdfDoc = await PDFDocument.load(pdfBytes); const pages = pdfDoc.getPages(); for (const page of pages) { const { width, height } = page.getSize(); // 方案1:替换为自定义归因文本(合规) page.drawText('翻译由Google提供支持', { x: width - 240, y: 22, size: 10, color: { r: 0.5, g: 0.5, b: 0.5 }, }); // 方案2:若需覆盖原标识,使用精准坐标的白色矩形 // page.drawRectangle({ // x: width - 250, // y: 20, // width: 240, // height: 20, // color: { r: 1, g: 1, b: 1 }, // }); } const modifiedPdfBytes = await pdfDoc.save(); fs.writeFileSync(`data/${crypto.randomUUID()}.pdf`, modifiedPdfBytes); }
说明
- 直接移除PDF中的文本并保留背景难度极大,因为PDF的文本与背景通常在同一图层,或文本为嵌入对象,没有原生API支持分离操作。
- 合规前提下,优先使用
customizedAttribution参数设置符合要求的自定义归因,避免后期处理。
内容的提问来源于stack exchange,提问作者Ajouve
相关产品推荐
相关产品推荐

