BeautifulSoup疑问:为何find_all传标签列表比逐个查询拼接更慢?
为什么多次调用
find_all单标签比一次调用多标签更快? 我发现一段多次调用find_all单标签的代码,运行速度远快于一次调用find_all传入标签列表的代码,这和我预期的相反——我原本以为前者要遍历5次文档树,后者只遍历一次,应该后者更快。
对比代码
更快的实现:
tags = ['a', 'h1', 'style', 'nav', 'footer'] result = [] for t in tags: for el in soup.find_all(t): result.append(el)
更慢的实现:
tags = ['a', 'h1', 'style', 'nav', 'footer'] result = [] for el in soup.find_all(tags): result.append(el)
计时代码
import timeit setup = ''' from bs4 import BeautifulSoup html = get_some_html_content() soup = BeautifulSoup(html) ''' stmt_1 = ''' result = [] for t in ['a', 'h1', 'style', 'nav', 'footer']: for el in soup.find_all(t): result.append(el) ''' stmt_2 = ''' result = [] for el in soup.find_all(['a', 'h1', 'style', 'nav', 'footer']): result.append(el) ''' print(timeit.timeit(setup=setup, stmt=stmt_1, number=100)) print(timeit.timeit(setup=setup, stmt=stmt_2, number=100))
原因分析
核心差异在于BeautifulSoup内部的查询逻辑:
- 传入单个标签字符串时,
find_all会直接调用内部的标签索引(类似哈希表结构),快速定位所有对应标签的元素,这个过程不需要遍历整个文档树,只是从索引中直接取出匹配元素,效率极高。 - 传入标签列表时,
find_all无法使用标签索引,只能遍历文档树中的每一个元素,再逐个检查元素标签是否在列表中。这个全文档遍历的开销,远大于5次单标签索引查询的总开销。
你的误解在于以为单标签查询会遍历文档树,但实际上BeautifulSoup在解析HTML时已经为标签建立了索引,单标签查询是直接走索引的,并没有全量遍历。
内容的提问来源于stack exchange,提问作者anon
相关产品推荐
相关产品推荐

