如何使用OpenXML逐行提取Azure Blob存储中Microsoft Word文档的文本
实现步骤及代码示例
前置依赖
先安装所需的NuGet包:
Azure.Storage.Blobs:用于操作Azure Blob存储DocumentFormat.OpenXml:用于处理Word docx格式文件
第一步:从Azure Blob读取文件到内存流
using Azure.Storage.Blobs; using System.IO; // 配置参数 string blobConnString = "你的Azure Blob连接字符串"; string containerName = "你的容器名称"; string blobFileName = "目标Word文件名称(需带.docx后缀)"; // 初始化Blob客户端 BlobContainerClient containerClient = new BlobContainerClient(blobConnString, containerName); BlobClient blobClient = containerClient.GetBlobClient(blobFileName); // 读取Blob到内存流 using MemoryStream memStream = new MemoryStream(); await blobClient.DownloadToAsync(memStream); // *关键:将流指针重置到起始位置,否则后续OpenXML会识别为文件损坏* memStream.Position = 0;
第二步:用OpenXML从内存流逐行提取文本
需要注意:Word文档本身没有严格的“行”定义,常规的行包括两种:段落分隔(对应OpenXML的w:p节点)、段落内的硬换行(对应w:br节点),以下代码兼容两种场景逐行输出:
using DocumentFormat.OpenXml.Packaging; using DocumentFormat.OpenXml.Wordprocessing; using System.Collections.Generic; List<string> allLines = new List<string>(); // 打开内存流中的Word文档,第二个参数设为false表示只读 using (WordprocessingDocument wordDoc = WordprocessingDocument.Open(memStream, false)) { Body body = wordDoc.MainDocumentPart.Document.Body; // 遍历所有段落 foreach (Paragraph para in body.Elements<Paragraph>()) { string currentLine = string.Empty; // 遍历段落内的所有元素 foreach (var element in para.Elements()) { if (element is Run run) { foreach (var runElement in run.Elements()) { // 遇到文本节点就追加到当前行 if (runElement is Text text) { currentLine += text.Text; } // 遇到硬换行就把当前行加入列表,重置当前行 else if (runElement is Break) { allLines.Add(currentLine); currentLine = string.Empty; } } } } // 段落结束,把剩余的当前行加入列表,不需要空行可以加判断过滤 allLines.Add(currentLine); } } // 最终allLines中存储的就是逐行提取的所有文本内容
常见注意事项
- OpenXML仅支持
.docx格式的Word文件,旧版.doc二进制格式无法使用该方法处理 - 如果不需要保留空行,在将
currentLine加入列表前可以加if (!string.IsNullOrWhiteSpace(currentLine))过滤 - 处理大文件时可以不用把所有行存入列表,每生成一行就直接处理,降低内存占用
内容的提问来源于stack exchange,提问作者David Herman
相关产品推荐
相关产品推荐

