1GB级200万行HTML表格快速筛选优化方案:高效提取Profession为Engineer的行
哥们,1GB的HTML表格+200万行数据,用BeautifulSoup确实会慢到离谱——毕竟它要把整个文档都加载到内存里构建DOM树,这光内存占用就够夸张的,更别说遍历修改的时间了。咱们换个流式解析的思路,不用把整个文件吃进内存,边读边处理,速度和内存占用都会友好太多!
为什么你的原代码这么慢?
原来的代码里fd.read()直接把1GB的文件全读进内存,然后BeautifulSoup要构建整个文档的DOM树,这一步就占了大量内存和时间;之后遍历所有<tr>再修改DOM,效率自然低得可怕。
更快的解决方案:流式解析
流式解析的核心是不加载整个文档到内存,逐行/逐元素处理,处理完就释放内存,适合超大规模文件。下面给你两种可行的方法:
方法1:用Python标准库HTMLParser(无需额外安装)
自定义一个解析器,只关注<tr>和<td>标签,只保留表头和职业为Engineer的行:
from html.parser import HTMLParser class EngineerTableParser(HTMLParser): def __init__(self): super().__init__() self.in_tr = False self.in_td = False self.current_td_index = 0 self.current_tr_data = [] self.output_lines = [] # 先写入HTML开头结构 self.output_lines.extend(["<html>", "<body>", "<table>"]) self.header_written = False def handle_starttag(self, tag, attrs): if tag == 'tr': self.in_tr = True self.current_tr_data = [] self.current_td_index = 0 # 记录<tr>的起始标签(包含属性) attr_str = ''.join([f' {key}="{value}"' for key, value in attrs]) self.current_tr_start = f"<tr{attr_str}>" elif tag == 'td' and self.in_tr: self.in_td = True def handle_endtag(self, tag): if tag == 'tr': self.in_tr = False # 处理表头或符合条件的行 if not self.header_written: self.output_lines.append(self.current_tr_start) for td_content in self.current_tr_data: self.output_lines.append(f"<td>{td_content}</td>") self.output_lines.append("</tr>") self.header_written = True elif self.current_tr_data[1].strip() == "Engineer": self.output_lines.append(self.current_tr_start) for td_content in self.current_tr_data: self.output_lines.append(f"<td>{td_content}</td>") self.output_lines.append("</tr>") elif tag == 'td' and self.in_tr: self.in_td = False self.current_td_index += 1 def handle_data(self, data): if self.in_td and self.in_tr: self.current_tr_data.append(data.strip()) def get_final_output(self): # 补充HTML结尾结构 self.output_lines.extend(["</table>", "</body>", "</html>"]) return '\n'.join(self.output_lines) # 处理文件 parser = EngineerTableParser() with open("input.html", 'r', encoding='utf-8') as input_file: for line in input_file: parser.feed(line) with open("output.html", 'w', encoding='utf-8') as output_file: output_file.write(parser.get_final_output())
方法2:用lxml(更快,推荐)
lxml是基于C实现的XML/HTML解析库,速度比标准库快很多,而且支持迭代解析,内存管理更高效。先安装:pip install lxml,然后用下面的代码:
from lxml import etree def filter_engineer_table(input_path, output_path): # 迭代解析<tr>标签,只处理start和end事件 context = etree.iterparse(input_path, events=('start', 'end'), tag='tr') output_content = [] # 写入HTML开头 output_content.extend(["<html>", "<body>", "<table>"]) header_written = False for event, elem in context: if event == 'end': if not header_written: # 写入表头行 output_content.append(etree.tostring(elem, encoding='unicode')) header_written = True else: # 获取第二列的文本内容 tds = elem.findall('td') if len(tds) >= 2 and tds[1].text.strip() == "Engineer": output_content.append(etree.tostring(elem, encoding='unicode')) # 清除当前元素,释放内存(关键!避免内存泄漏) elem.clear() # 清理父节点的引用 while elem.getprevious() is not None: del elem.getparent()[0] # 写入HTML结尾 output_content.extend(["</table>", "</body>", "</html>"]) with open(output_path, 'w', encoding='utf-8') as output_file: output_file.write('\n'.join(output_content)) # 调用函数处理文件 filter_engineer_table("input.html", "output.html")
效果对比
这两种方法都是边读边处理,内存占用只有几MB到几十MB,完全不会像原代码那样把1GB文件加载到内存里。其中lxml的方法速度最快,处理1GB的文件应该能把时间从15小时压缩到几十分钟(具体看你的硬件,但肯定是数量级的提升)。
内容的提问来源于stack exchange,提问作者PJ47
相关产品推荐
相关产品推荐

