如何优化嵌套循环与DataFrame创建以提升HTML解析脚本性能?
问题:提升HTML解析脚本的运行效率
我编程经验尚浅,正在开发一个基于customtkinter的脚本,用户可输入包含诊断地址及对应信息的特定HTML文件,脚本会解析该文件并返回选中的地址/信息字典供后续使用。
目前脚本可正常运行,但处理的HTML文件行数在10000至70000行不等,较大文件的解析耗时超过1分钟。我知道代码存在效率问题,正尝试寻找缩短运行等待时间的方法,认为重复创建DataFrame及后续的嵌套循环是主要性能瓶颈,曾考虑只创建一个DataFrame并遍历,但不确定效果。
请问如何编写代码以提升运行效率?
原相关函数代码
# Clear list of previous selections fv.info_values_to_add.clear() # Read user's info selections filter_info_selections() # Open and read protocol file with open(file_name, 'r') as file: contents = file.read() # Global variable to used in export functions global Length_Of_Info_1 Length_Of_Info_1 = len(fv.info_values_to_add) # Create a soup object from the protocol parsed_protocol = BeautifulSoup(contents, "html.parser") for address, address_value in fv.protocol_values_1.items(): # String to be used to find the correct section of the html string_address = "ECU: " + (address) try: # Find the header for the parsed address table = parsed_protocol.find('p', string = re.compile(string_address)) # Select the correct table for the information data_table = table.find_all_next('table') # Create a dataframe from the table data_frame = pd.read_html(io.StringIO(str(data_table)))[1] # Clean data frame columns and values df_clean = data_frame.drop(columns=2, axis=1) # Save selected data to variables to be used sw_version = df_clean.iloc[1,1] hw_part_number = df_clean.iloc[2,1] hw_version = df_clean.iloc[3,1] vehicle_vin = df_clean.iloc[20,1] fazit_id = df_clean.iloc[21,1] coding = df_clean.iloc[7,1] vw_part_number = df_clean.iloc[0,1] # List to store variables to be added to the fv.protocol_values_1 dictionary temp_list = [] # Iterate through the info list and add the selected variables for key in fv.info_values_to_add: if key == "Software Version": temp_list.append(sw_version) elif key == "Hardware part number": temp_list.append(hw_part_number) elif key == "Hardware Version": temp_list.append(hw_version) elif key == "Fazit ID": temp_list.append(fazit_id) elif key == "VIN Number": temp_list.append(vehicle_vin) elif key == "Coding": temp_list.append(coding) elif key == "VW part number": temp_list.append(vw_part_number) else: pass # Add values to the address in the dictionary fv.protocol_values_1[address] = temp_list
优化方案及代码实现
核心优化点
- 更换解析器:用
lxml替代默认的html.parser,解析速度提升明显(需先安装lxml库:pip install lxml) - 跳过DataFrame:直接用BeautifulSoup提取表格指定位置的内容,避免重复创建DataFrame的开销
- 字典映射替代条件判断:把字段名和对应表格位置的映射提前定义,循环时直接取值,减少嵌套判断
- 简化匹配逻辑:用字符串包含判断替代重复编译正则表达式,降低匹配开销
优化后的代码
# 确保已安装lxml库 from bs4 import BeautifulSoup # Clear list of previous selections fv.info_values_to_add.clear() # Read user's info selections filter_info_selections() # Open and read protocol file with open(file_name, 'r') as file: contents = file.read() # Global variable to used in export functions global Length_Of_Info_1 Length_Of_Info_1 = len(fv.info_values_to_add) # 使用lxml解析器,速度远快于默认html.parser parsed_protocol = BeautifulSoup(contents, "lxml") # 预定义字段与表格行索引的映射(对应原df_clean.iloc[row,1]的row值) field_row_map = { "Software Version": 1, "Hardware part number": 2, "Hardware Version": 3, "Fazit ID": 21, "VIN Number": 20, "Coding": 7, "VW part number": 0 } for address, _ in fv.protocol_values_1.items(): string_address = f"ECU: {address}" try: # 用字符串包含匹配替代正则,减少编译开销 p_tag = parsed_protocol.find('p', string=lambda text: text and string_address in text) if not p_tag: continue # 获取目标表格(原代码取第二个表格,即索引1) data_tables = p_tag.find_all_next('table') if len(data_tables) < 2: continue target_table = data_tables[1] # 提取表格所有行的第二列值,跳过原列2的内容 table_values = [] rows = target_table.find_all('tr') for row in rows: tds = row.find_all('td') if len(tds) >= 2: table_values.append(tds[1].get_text(strip=True)) # 根据用户选择的字段生成结果列表 temp_list = [] for key in fv.info_values_to_add: row_idx = field_row_map.get(key) if row_idx is not None and row_idx < len(table_values): temp_list.append(table_values[row_idx]) else: temp_list.append("") # 无对应值时填充空字符串 # 更新字典值 fv.protocol_values_1[address] = temp_list except Exception as e: # 单个地址解析失败不中断整体流程,可添加日志记录 print(f"解析地址{address}时出错: {str(e)}") fv.protocol_values_1[address] = []
额外优化建议
- 批量提取ECU信息:如果HTML中ECU地址的
<p>标签格式统一,可一次性提取所有ECU地址及对应表格,再与fv.protocol_values_1中的地址匹配,减少多次find操作的开销 - 减少全局变量依赖:尽量通过函数返回值传递
Length_Of_Info_1,避免使用全局变量 - 多进程处理:若解析任务为CPU密集型,可拆分任务用多进程处理(规避Python GIL限制)
内容的提问来源于stack exchange,提问作者Chickchu
相关产品推荐
相关产品推荐

