You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化嵌套循环与DataFrame创建以提升HTML解析脚本性能?

问题:提升HTML解析脚本的运行效率

我编程经验尚浅,正在开发一个基于customtkinter的脚本,用户可输入包含诊断地址及对应信息的特定HTML文件,脚本会解析该文件并返回选中的地址/信息字典供后续使用。

目前脚本可正常运行,但处理的HTML文件行数在10000至70000行不等,较大文件的解析耗时超过1分钟。我知道代码存在效率问题,正尝试寻找缩短运行等待时间的方法,认为重复创建DataFrame及后续的嵌套循环是主要性能瓶颈,曾考虑只创建一个DataFrame并遍历,但不确定效果。

请问如何编写代码以提升运行效率?

原相关函数代码

# Clear list of previous selections
fv.info_values_to_add.clear()

# Read user's info selections
filter_info_selections()

# Open and read protocol file
with open(file_name, 'r') as file:
    contents = file.read()

# Global variable to used in export functions
global Length_Of_Info_1
Length_Of_Info_1 = len(fv.info_values_to_add)

# Create a soup object from the protocol
parsed_protocol = BeautifulSoup(contents, "html.parser")


for address, address_value in fv.protocol_values_1.items():


    # String to be used to find the correct section of the html
    string_address = "ECU: " + (address)

    try:
        
        # Find the header for the parsed address
        table = parsed_protocol.find('p', string = re.compile(string_address))
        
        # Select the correct table for the information
        data_table = table.find_all_next('table')

        # Create a dataframe from the table
        data_frame = pd.read_html(io.StringIO(str(data_table)))[1]

        # Clean data frame columns and values
        df_clean = data_frame.drop(columns=2, axis=1)
        
        # Save selected data to variables to be used
        sw_version = df_clean.iloc[1,1]
        hw_part_number = df_clean.iloc[2,1]
        hw_version = df_clean.iloc[3,1]
        vehicle_vin = df_clean.iloc[20,1]
        fazit_id = df_clean.iloc[21,1]
        coding = df_clean.iloc[7,1]
        vw_part_number = df_clean.iloc[0,1]


        # List to store variables to be added to the fv.protocol_values_1 dictionary
        temp_list = []

        
        # Iterate through the info list and add the selected variables
        for key in fv.info_values_to_add:

            if key == "Software Version":
                temp_list.append(sw_version)
            

            elif key == "Hardware part number":
                temp_list.append(hw_part_number)
            

            elif key == "Hardware Version":
                temp_list.append(hw_version)
            

            elif key == "Fazit ID":
                temp_list.append(fazit_id)
            

            elif key == "VIN Number":
                temp_list.append(vehicle_vin)
            

            elif key == "Coding":
                temp_list.append(coding)
            

            elif key == "VW part number":
                temp_list.append(vw_part_number)


            else:
                pass
            
                
        # Add values to the address in the dictionary
        fv.protocol_values_1[address] = temp_list

优化方案及代码实现

核心优化点

  • 更换解析器:用lxml替代默认的html.parser,解析速度提升明显(需先安装lxml库:pip install lxml)
  • 跳过DataFrame:直接用BeautifulSoup提取表格指定位置的内容,避免重复创建DataFrame的开销
  • 字典映射替代条件判断:把字段名和对应表格位置的映射提前定义,循环时直接取值,减少嵌套判断
  • 简化匹配逻辑:用字符串包含判断替代重复编译正则表达式,降低匹配开销

优化后的代码

# 确保已安装lxml库
from bs4 import BeautifulSoup

# Clear list of previous selections
fv.info_values_to_add.clear()

# Read user's info selections
filter_info_selections()

# Open and read protocol file
with open(file_name, 'r') as file:
    contents = file.read()

# Global variable to used in export functions
global Length_Of_Info_1
Length_Of_Info_1 = len(fv.info_values_to_add)

# 使用lxml解析器,速度远快于默认html.parser
parsed_protocol = BeautifulSoup(contents, "lxml")

# 预定义字段与表格行索引的映射(对应原df_clean.iloc[row,1]的row值)
field_row_map = {
    "Software Version": 1,
    "Hardware part number": 2,
    "Hardware Version": 3,
    "Fazit ID": 21,
    "VIN Number": 20,
    "Coding": 7,
    "VW part number": 0
}

for address, _ in fv.protocol_values_1.items():
    string_address = f"ECU: {address}"
    
    try:
        # 用字符串包含匹配替代正则,减少编译开销
        p_tag = parsed_protocol.find('p', string=lambda text: text and string_address in text)
        if not p_tag:
            continue
            
        # 获取目标表格(原代码取第二个表格,即索引1)
        data_tables = p_tag.find_all_next('table')
        if len(data_tables) < 2:
            continue
        target_table = data_tables[1]
        
        # 提取表格所有行的第二列值,跳过原列2的内容
        table_values = []
        rows = target_table.find_all('tr')
        for row in rows:
            tds = row.find_all('td')
            if len(tds) >= 2:
                table_values.append(tds[1].get_text(strip=True))
        
        # 根据用户选择的字段生成结果列表
        temp_list = []
        for key in fv.info_values_to_add:
            row_idx = field_row_map.get(key)
            if row_idx is not None and row_idx < len(table_values):
                temp_list.append(table_values[row_idx])
            else:
                temp_list.append("")  # 无对应值时填充空字符串
        
        # 更新字典值
        fv.protocol_values_1[address] = temp_list
        
    except Exception as e:
        # 单个地址解析失败不中断整体流程,可添加日志记录
        print(f"解析地址{address}时出错: {str(e)}")
        fv.protocol_values_1[address] = []

额外优化建议

  • 批量提取ECU信息:如果HTML中ECU地址的<p>标签格式统一,可一次性提取所有ECU地址及对应表格,再与fv.protocol_values_1中的地址匹配,减少多次find操作的开销
  • 减少全局变量依赖:尽量通过函数返回值传递Length_Of_Info_1,避免使用全局变量
  • 多进程处理:若解析任务为CPU密集型,可拆分任务用多进程处理(规避Python GIL限制)

内容的提问来源于stack exchange,提问作者Chickchu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 09:37:38