You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

1GB级200万行HTML表格快速筛选优化方案:高效提取Profession为Engineer的行

哥们,1GB的HTML表格+200万行数据,用BeautifulSoup确实会慢到离谱——毕竟它要把整个文档都加载到内存里构建DOM树,这光内存占用就够夸张的,更别说遍历修改的时间了。咱们换个流式解析的思路,不用把整个文件吃进内存,边读边处理,速度和内存占用都会友好太多!

为什么你的原代码这么慢?

原来的代码里fd.read()直接把1GB的文件全读进内存,然后BeautifulSoup要构建整个文档的DOM树,这一步就占了大量内存和时间;之后遍历所有<tr>再修改DOM,效率自然低得可怕。

更快的解决方案:流式解析

流式解析的核心是不加载整个文档到内存,逐行/逐元素处理,处理完就释放内存,适合超大规模文件。下面给你两种可行的方法:

方法1:用Python标准库HTMLParser(无需额外安装)

自定义一个解析器,只关注<tr>和<td>标签,只保留表头和职业为Engineer的行:

from html.parser import HTMLParser

class EngineerTableParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_tr = False
        self.in_td = False
        self.current_td_index = 0
        self.current_tr_data = []
        self.output_lines = []
        # 先写入HTML开头结构
        self.output_lines.extend(["<html>", "<body>", "<table>"])
        self.header_written = False

    def handle_starttag(self, tag, attrs):
        if tag == 'tr':
            self.in_tr = True
            self.current_tr_data = []
            self.current_td_index = 0
            # 记录<tr>的起始标签(包含属性)
            attr_str = ''.join([f' {key}="{value}"' for key, value in attrs])
            self.current_tr_start = f"<tr{attr_str}>"
        elif tag == 'td' and self.in_tr:
            self.in_td = True

    def handle_endtag(self, tag):
        if tag == 'tr':
            self.in_tr = False
            # 处理表头或符合条件的行
            if not self.header_written:
                self.output_lines.append(self.current_tr_start)
                for td_content in self.current_tr_data:
                    self.output_lines.append(f"<td>{td_content}</td>")
                self.output_lines.append("</tr>")
                self.header_written = True
            elif self.current_tr_data[1].strip() == "Engineer":
                self.output_lines.append(self.current_tr_start)
                for td_content in self.current_tr_data:
                    self.output_lines.append(f"<td>{td_content}</td>")
                self.output_lines.append("</tr>")
        elif tag == 'td' and self.in_tr:
            self.in_td = False
            self.current_td_index += 1

    def handle_data(self, data):
        if self.in_td and self.in_tr:
            self.current_tr_data.append(data.strip())

    def get_final_output(self):
        # 补充HTML结尾结构
        self.output_lines.extend(["</table>", "</body>", "</html>"])
        return '\n'.join(self.output_lines)

# 处理文件
parser = EngineerTableParser()
with open("input.html", 'r', encoding='utf-8') as input_file:
    for line in input_file:
        parser.feed(line)

with open("output.html", 'w', encoding='utf-8') as output_file:
    output_file.write(parser.get_final_output())

方法2:用lxml(更快,推荐)

lxml是基于C实现的XML/HTML解析库,速度比标准库快很多,而且支持迭代解析,内存管理更高效。先安装:pip install lxml,然后用下面的代码:

from lxml import etree

def filter_engineer_table(input_path, output_path):
    # 迭代解析<tr>标签,只处理start和end事件
    context = etree.iterparse(input_path, events=('start', 'end'), tag='tr')
    
    output_content = []
    # 写入HTML开头
    output_content.extend(["<html>", "<body>", "<table>"])
    header_written = False

    for event, elem in context:
        if event == 'end':
            if not header_written:
                # 写入表头行
                output_content.append(etree.tostring(elem, encoding='unicode'))
                header_written = True
            else:
                # 获取第二列的文本内容
                tds = elem.findall('td')
                if len(tds) >= 2 and tds[1].text.strip() == "Engineer":
                    output_content.append(etree.tostring(elem, encoding='unicode'))
            # 清除当前元素,释放内存(关键!避免内存泄漏)
            elem.clear()
            # 清理父节点的引用
            while elem.getprevious() is not None:
                del elem.getparent()[0]

    # 写入HTML结尾
    output_content.extend(["</table>", "</body>", "</html>"])

    with open(output_path, 'w', encoding='utf-8') as output_file:
        output_file.write('\n'.join(output_content))

# 调用函数处理文件
filter_engineer_table("input.html", "output.html")

效果对比

这两种方法都是边读边处理,内存占用只有几MB到几十MB,完全不会像原代码那样把1GB文件加载到内存里。其中lxml的方法速度最快,处理1GB的文件应该能把时间从15小时压缩到几十分钟(具体看你的硬件,但肯定是数量级的提升)。


内容的提问来源于stack exchange,提问作者PJ47

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 05:14:05