如何从Word文档中提取带文件名的嵌入式纯文本文件?
提取Docx中嵌入式纯文本文件的可行方案
方法一:Python脚本解析(推荐,无需VBA)
Docx本质是ZIP压缩包,其中word/embeddings/下的.bin文件是OLE格式容器,包含嵌入的文本内容和原文件名。可以通过olefile库解析这些容器:
步骤:
- 安装依赖库:
pip install olefile
- 运行以下脚本:
import zipfile import olefile import os def extract_embedded_text(docx_path, output_dir): os.makedirs(output_dir, exist_ok=True) with zipfile.ZipFile(docx_path, 'r') as zf: for entry in zf.infolist(): if entry.filename.startswith('word/embeddings/') and entry.filename.endswith('.bin'): with zf.open(entry) as f: bin_data = f.read() if olefile.isOleFile(bin_data): with olefile.OleFileIO(bin_data) as ole: for stream in ole.listdir(): if stream[0] == 'Package': # 获取原文件名,失败则用bin文件名替换后缀 try: filename = ole.get_metadata().filename except: filename = os.path.basename(entry.filename).replace('.bin', '.txt') text_content = ole.openstream(stream).read().decode('utf-8', errors='ignore') output_path = os.path.join(output_dir, filename) with open(output_path, 'w', encoding='utf-8') as out_f: out_f.write(text_content) print(f"已提取:{filename}") # 替换为你的docx路径和输出目录 extract_embedded_text("你的文档.docx", "提取结果")
方法二:LibreOffice GUI操作(无需编程)
适合不想写代码的场景:
- 用LibreOffice Writer打开目标docx文件
- 逐个选中嵌入式文本文件,右键选择保存副本,系统会自动保留原文件名
- 若嵌入文件较多,可使用LibreOffice宏批量处理(但效率不如Python脚本)
方法三:改进POI的Java实现
如果需要在Java环境下处理,需用POIFSFileSystem解析OLE格式的.bin文件,而不是直接读取ZipPackagePart内容:
import org.apache.poi.poifs.filesystem.POIFSFileSystem; import org.apache.poi.openxml4j.opc.OPCPackage; import org.apache.poi.openxml4j.opc.PackagePart; import java.io.*; import java.util.List; public class ExtractEmbeddedTxt { public static void main(String[] args) throws Exception { String docxPath = "你的文档.docx"; String outputDir = "提取结果"; new File(outputDir).mkdirs(); try (OPCPackage pkg = OPCPackage.open(new File(docxPath))) { List<PackagePart> embedParts = pkg.getPartsByNameRegex("/word/embeddings/.*\\.bin"); for (PackagePart part : embedParts) { try (InputStream is = part.getInputStream(); POIFSFileSystem poifs = new POIFSFileSystem(is)) { // 获取原文件名,为空则用bin文件名替代 String filename = null; if (poifs.getRoot().getDocumentSummaryInformation() != null) { filename = poifs.getRoot().getDocumentSummaryInformation().getFileName(); } if (filename == null || filename.isEmpty()) { filename = part.getPartName().getName().substring(part.getPartName().getName().lastIndexOf('/') + 1).replace(".bin", ".txt"); } try (InputStream txtIs = poifs.createDocumentInputStream("Package"); BufferedReader br = new BufferedReader(new InputStreamReader(txtIs, "UTF-8")); BufferedWriter bw = new BufferedWriter(new FileWriter(new File(outputDir, filename)))) { String line; while ((line = br.readLine()) != null) { bw.write(line); bw.newLine(); } System.out.println("已提取:" + filename); } } } } } }
内容的提问来源于stack exchange,提问作者Liam lin
相关产品推荐
相关产品推荐

