如何用Python提取.docm文件表格信息?遇阻求替代路径
Hey, let's tackle this problem head-on—you’ve got 700+ .docm files with embedded tables, python-docx throws errors when you try to work with them, and you don’t have time to learn VB or C#? No sweat, here are two practical, no-fuss solutions that fit your needs:
1. 批量转.docx用LibreOffice(极简操作,几乎无代码)
If you want to stick with python-docx and just need to get past the .docm barrier, LibreOffice’s command-line tool lets you batch-convert all your files to .docx in one go—no VB/C# required. Here’s how:
- First, install LibreOffice (it’s free, open-source, and works on Windows/Mac/Linux)
- Open a command prompt (Windows) or terminal (Mac/Linux), navigate to the folder where your .docm files are stored
- Run this single command to convert every .docm file in the folder:
soffice --headless --convert-to docx *.docm
Once the conversion finishes, you can use python-docx exactly as you planned to extract table data from the new .docx files.
2. 直接用Python提取.docm表格(无需转换文件)
If you’re on a Windows machine with Microsoft Word installed, you can skip conversion entirely and pull table data directly from .docm files using the pywin32 library. This uses Word’s built-in COM interface, but you don’t need to know VB—just a few lines of Python:
- First, install the dependency:
pip install pywin32 - Use this sample code to batch-extract tables and save them as CSV files (easy to work with later):
import win32com.client as win32 import os import csv def extract_tables_from_docm(docm_path, output_dir): # Launch Word in background mode word = win32.gencache.EnsureDispatch('Word.Application') word.Visible = False try: # Open the .docm file doc = word.Documents.Open(os.path.abspath(docm_path)) file_base_name = os.path.splitext(os.path.basename(docm_path))[0] # Loop through each table in the document for table_num, table in enumerate(doc.Tables, 1): csv_file_path = os.path.join(output_dir, f"{file_base_name}_table_{table_num}.csv") with open(csv_file_path, 'w', encoding='utf-8', newline='') as csv_file: writer = csv.writer(csv_file) # Extract each row's cell text for row in table.Rows: cell_contents = [cell.Range.Text.strip().replace('\r', '') for cell in row.Cells] writer.writerow(cell_contents) print(f"Success: Extracted {len(doc.Tables)} tables from {docm_path}") except Exception as e: print(f"Failed to process {docm_path}: {str(e)}") finally: # Clean up to avoid leaving Word running in background if 'doc' in locals(): doc.Close(SaveChanges=False) word.Quit() # Configure your input/output folders input_folder = "./your_docm_files" # Replace with your folder path output_folder = "./extracted_table_data" os.makedirs(output_folder, exist_ok=True) # Process all .docm files in the input folder for filename in os.listdir(input_folder): if filename.lower().endswith('.docm'): extract_tables_from_docm(os.path.join(input_folder, filename), output_folder)
This code will save each table from every .docm as a separate CSV file, giving you clean, structured data without ever converting the original files. Just make sure Word isn’t open manually while running the script—this avoids conflicts with the background Word instance.
内容的提问来源于stack exchange,提问作者James Cooper

