Python解析大型XML文件时如何添加计数逻辑统计指定匹配结果数量
代码改造方案
计数逻辑位置说明
计数逻辑直接放在现有内层循环的匹配分支中最优,无需二次遍历XML节点,和原有提取逻辑复用同一次遍历过程,不会额外增加性能开销,适合你当前1500条以上记录的大文件场景。
最优实现方案(基础计数器版本)
你只需要统计两个固定值的出现次数,直接用两个整型变量计数即可,逻辑最简洁,无额外依赖:
- 在遍历节点的循环前初始化两个计数器
- 每次匹配到目标值时给对应计数器+1
- 写入HTML的末尾输出统计结果
完整改造后代码如下:
import webbrowser from lxml import etree tree = etree.parse('g_d_anime.xml') root = tree.getroot() tree.findall('WorksXML/Work') h_file = open("xml_stats.html","w") h_file.write("This block shows the producers and genres for this file.") h_file.write('<br>') # 初始化计数器 aniplex_count = 0 magic_count = 0 # 处理制片方 for Producers in root.iter('Producers'): p = Producers.text.split(',') for producer in p: if producer == 'Aniplex': print(p) h_file.write('<li>' + str(producer) + '</li><br>') aniplex_count += 1 # 匹配到就计数 # 处理流派 for Genres in root.iter('Genres'): g = Genres.text.split(',') for genre in g: if genre == 'Magic': print(g) h_file.write('<li>'+ str(genre) + '</li><br>') magic_count += 1 # 匹配到就计数 # 写入统计结果 stat_text = f"Aniplex: {aniplex_count}、Magic: {magic_count}" print(stat_text) h_file.write(f"<h3>统计结果:{stat_text}</h3>") h_file.close() webbrowser.open_new_tab("xml_stats.html")
扩展方案(Counter版本,适合多值统计)
如果你之后需要统计所有制片方、流派的出现次数,可以用collections.Counter实现,代码示例如下:
from collections import Counter # 初始化计数器 producer_counter = Counter() genre_counter = Counter() for Producers in root.iter('Producers'): p_list = Producers.text.split(',') producer_counter.update(p_list) for Genres in root.iter('Genres'): g_list = Genres.text.split(',') genre_counter.update(g_list) # 提取目标值 aniplex_count = producer_counter.get('Aniplex', 0) magic_count = genre_counter.get('Magic', 0)
内容的提问来源于stack exchange,提问作者SassyG
相关产品推荐
相关产品推荐

