如何使用python-docx库准确统计Word文档页数?无需转PDF
用python-docx统计Word文档页数(基于页脚页码)
由于python-docx没有原生API直接获取Word文档的页数(Word页数是渲染时动态计算的),但结合你提到的「页脚页码为每页唯一内容」的条件,可以通过提取所有页脚中的页码并去重计数的方式实现准确统计,替代不靠谱的「统计节」方法。
实现步骤与代码
from docx import Document import re def count_pages_via_footer(doc_path): doc = Document(doc_path) page_numbers = set() # 遍历所有节 for section in doc.sections: # 遍历该节的所有页脚类型(普通页、首页、奇偶页) for footer_type in [section.footer, section.first_page_footer, section.even_page_footer]: if not footer_type.is_linked_to_previous: # 提取页脚所有段落的文本 footer_text = "\n".join([para.text.strip() for para in footer_type.paragraphs]) # 匹配页码(假设页码为数字,可根据实际格式调整正则) matches = re.findall(r"\d+", footer_text) for match in matches: page_numbers.add(int(match)) # 集合的长度即为总页数 return len(page_numbers) # 使用示例 doc_path = "your_document.docx" total_pages = count_pages_via_footer(doc_path) print(f"文档总页数:{total_pages}")
注意事项
- 如果文档使用罗马数字页码(如Ⅰ、Ⅱ或i、ii),需要修改正则表达式或添加数字转换逻辑,比如用
roman库将罗马数字转为阿拉伯数字后再加入集合。 - 若部分节的页脚与上一节链接(
is_linked_to_previous=True),会跳过重复的页脚内容,避免重复统计页码。 - 确保页脚中只有页码内容,若有其他文本,需调整正则表达式精准匹配页码(比如匹配位于页脚特定位置的数字)。
内容的提问来源于stack exchange,提问作者Vedant Bhatt
相关产品推荐
相关产品推荐

