如何通过Power Automate或编程实现SharePoint Online Word文件版本对比(开源)
方案1:Power Automate 结合开源文本对比工具
步骤拆解
获取文件版本内容
用Power Automate内置的SharePoint Online连接器完成:- 调用
Get file content获取当前版本文件流; - 调用
Get file version history拿到目标历史版本的ID,再通过Get file content using path指定版本ID提取历史版本文件流; - 将Word文件转纯文本:用开源工具
pandoc(可通过Power Automate的Run script动作执行命令:pandoc -s "input.docx" -o "output.txt")。
- 调用
生成Git Diff样式对比结果
用Google开源的diff-match-patch库,通过Power Automate的Run JavaScript code动作调用,输出Git风格的差异内容:const DiffMatchPatch = require('diff-match-patch'); const dmp = new DiffMatchPatch(); const diffs = dmp.diff_main(currentText, historyText); dmp.diff_cleanupSemantic(diffs); let gitDiff = '--- Current Version\n+++ Historical Version\n'; diffs.forEach(diff => { switch(diff[0]) { case 1: gitDiff += `+${diff[1]}\n`; break; case -1: gitDiff += `-${diff[1]}\n`; break; case 0: gitDiff += ` ${diff[1]}\n`; break; } }); return gitDiff;
方案2:Python 编程实现
依赖库
shareplum(操作SharePoint)、python-docx(读取Word内容)、difflib(生成diff)
核心代码
import difflib from docx import Document from shareplum import Site # 连接SharePoint Online site = Site('https://yourtenant.sharepoint.com/sites/yoursite', username='user@tenant.com', password='your_password') folder = site.Folder('Shared Documents/TargetFolder') # 获取当前版本与历史版本文件流 current_file = folder.get_file('TargetFile.docx') versions = folder.get_file_versions('TargetFile.docx') historical_file = folder.get_file_version('TargetFile.docx', version_id=versions[0]['ID']) # 提取Word纯文本 def extract_docx_text(file_stream): doc = Document(file_stream) return '\n'.join([para.text for para in doc.paragraphs]) current_text = extract_docx_text(current_file) history_text = extract_docx_text(historical_file) # 生成Git Diff样式输出 diff_result = difflib.unified_diff( current_text.splitlines(), history_text.splitlines(), fromfile='Current Version', tofile='Historical Version', lineterm='' ) print('\n'.join(diff_result))
方案3:.NET 编程实现
依赖库
Microsoft.SharePointOnline.CSOM(操作SharePoint)、DocumentFormat.OpenXml(读取Word)、DiffPlex(生成diff)
核心代码片段
using DiffPlex; using DiffPlex.Model; using Microsoft.SharePoint.Client; using DocumentFormat.OpenXml.Packaging; // 初始化SharePoint连接 ClientContext ctx = new ClientContext("https://yourtenant.sharepoint.com/sites/yoursite"); ctx.Credentials = new SharePointOnlineCredentials("user@tenant.com", GetSecurePassword()); File targetFile = ctx.Web.GetFileByServerRelativeUrl("/sites/yoursite/Shared Documents/TargetFile.docx"); ctx.Load(targetFile, f => f.Versions); ctx.ExecuteQuery(); // 提取两个版本的文本内容 string currentText = ExtractWordText(targetFile.OpenBinaryStream()); FileVersion historicalVersion = targetFile.Versions[0]; string historyText = ExtractWordText(historicalVersion.OpenBinaryStream()); // 生成Git风格差异 var differ = new Differ(); DiffResult diffModel = differ.CreateDiffModel(currentText, historyText); StringBuilder gitDiff = new StringBuilder("--- Current Version\n+++ Historical Version\n"); foreach (var line in diffModel.Lines) { switch (line.Type) { case ChangeType.Inserted: gitDiff.AppendLine($"+{line.Text}"); break; case ChangeType.Deleted: gitDiff.AppendLine($"-{line.Text}"); break; case ChangeType.Unchanged: gitDiff.AppendLine($" {line.Text}"); break; } } // 输出或保存结果 Console.WriteLine(gitDiff.ToString()); // 辅助方法:提取Word文本 private static string ExtractWordText(System.IO.Stream stream) { using (WordprocessingDocument doc = WordprocessingDocument.Open(stream, false)) { return doc.MainDocumentPart.Document.Body.InnerText; } }
注意事项
- Word中的表格、图片等非文本内容,开源工具无法完全提取,可根据需求标记为
[表格]/[图片]或忽略; - Power Automate中运行脚本时,可通过Azure Function托管代码,或使用Power Automate的Python脚本预览功能(需配置Python环境)。
内容的提问来源于stack exchange,提问作者Ramesh
相关产品推荐
相关产品推荐

