如何统计.odt文档单词数?遍历文件夹批量实现方案
批量统计ODT文档单词数的实现需求与代码
功能需求
- 遍历指定文件夹及其子文件夹,筛选所有OpenOffice
.odt文档 - 打开每个
.odt文档并统计其中的单词数量 - 汇总所有文档的单词数并输出结果
目前使用odfpy库操作.odt文件,但现有示例和文档多聚焦于内容添加、样式获取等操作,缺少提取文档文本的相关资料。
单文件单词统计代码
基于字符统计思路调整出单文件单词统计代码(当前仅处理指定test.odt文件,后续需优化为批量处理):
def count_words_in_file(file_list): # 打开所有找到的.odt文件,统计单词数并求和 # 后续需调整为处理所有文件 file_path = "test.odt" from odf import text # 读取文档 document_text = load(file_path) # 获取文档中所有段落 all_paragraphs = document_text.getElementsByType(text.P) final_word_count = 0 # 遍历每个段落,提取文本并统计单词数 for paragraph in all_paragraphs: text_content = teletype.extractText(paragraph) words = text_content.split(" ") while '' in words: words.remove('') print(words) final_word_count += len(words) print(f"最终单词数: {final_word_count}")
文件夹遍历初始代码
已编写文件夹遍历的初始代码,目前需完善批量统计逻辑:
# 统计指定文件夹及其子文件夹中的.odt文档数量与总单词数 # 默认检查脚本所在目录的上级目录 # 导入所需库 import os from odf.opendocument import OpenDocumentText from odf import text def main(): # 获取脚本当前所在目录 current_dir = os.path.dirname(os.path.abspath(__file__)) # 设置为上级目录 above_dir = current_dir + "\.." # 调用函数扫描.odt文件 file_list = scan_for_files(above_dir) # 调用函数统计单词数 count_words_in_file(file_list) def scan_for_files(above_dir): # 存储所有找到的文件路径 file_list = [] # 遍历所有文件夹及子文件夹 for folder, subfolder, files in os.walk(above_dir): for file in files: complete_path = os.path.join(folder, file) file_list.append(complete_path) return file_list def count_words_in_file(file_list): # 打开所有找到的.odt文件,统计单词数并求和 for file in file_list: if file.endswith(".odt"): textdoc = OpenDocumentText() for paragraph in textdoc.body.childNodes: print(paragraph) main()
内容的提问来源于stack exchange,提问作者Arvid Eriksson
相关产品推荐
相关产品推荐

