You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在BeautifulSoup的find_all方法中使用正则表达式实现带优先级的匹配

实现带优先级的标签匹配(按顺序匹配,匹配到即停止后续模式)

你遇到的问题是因为正则表达式的|是逻辑或,只要标签名符合任意一个模式就会被选中,而BeautifulSoup的find_all会遍历所有标签并收集所有匹配项,不会自动帮你实现“优先匹配第一个模式,成功就忽略后面”的逻辑。

下面给你两种可行的解决方案,都能实现你要的优先级匹配效果:

方案一:分步查找(最直观易维护)

按优先级顺序依次查找,找到第一个有结果的模式就停止,直接返回该结果:

from bs4 import BeautifulSoup
import re

xml_doc = """
<m3_commodity_group commodity3="Oilseeds"><m3_year_group_Collection><m3_year_group market_year3="2011/12"><m3_month_group_Collection><m3_month_group forecast_month3=""><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell cell_value3="353.93"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell cell_value3="429.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell cell_value3="73.59"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell cell_value3="345.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell cell_value3="59.03"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection><m3_value_group_Collection><m3_value_group><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection></m3_value_group></m3_value_group_Collection></m3_month_group></m3_month_group_Collection></m3_year_group></m3_year_group_Collection></m3_commodity_group>
"""
soup = BeautifulSoup(xml_doc, "xml")

# 定义优先级模式列表,顺序就是优先级
priority_patterns = [
    r'^m[0-9]_commodity_group$',
    r'^m[0-9]_region_group$',
    r'^m[0-9]_attribute_group$'
]

matched_results = []
for pattern in priority_patterns:
    current_matches = soup.find_all(re.compile(pattern, flags=re.I))
    if current_matches:
        matched_results = current_matches
        break  # 找到第一个匹配的模式,直接停止后续查找

print(len(matched_results))  # 输出1,和你预期的一致

方案二:自定义匹配函数(适合一次性传递逻辑)

利用BeautifulSoup支持传入自定义函数作为匹配条件的特性,在函数内部实现优先级判断:

from bs4 import BeautifulSoup
import re

xml_doc = """
<m3_commodity_group commodity3="Oilseeds"><m3_year_group_Collection><m3_year_group market_year3="2011/12"><m3_month_group_Collection><m3_month_group forecast_month3=""><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell cell_value3="353.93"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell cell_value3="429.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell cell_value3="73.59"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell cell_value3="345.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell cell_value3="59.03"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection><m3_value_group_Collection><m3_value_group><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection></m3_value_group></m3_value_group_Collection></m3_month_group></m3_month_group_Collection></m3_year_group></m3_year_group_Collection></m3_commodity_group>
"""
soup = BeautifulSoup(xml_doc, "xml")

def priority_tag_match(tag):
    # 按优先级检查:先看是否匹配第一个模式
    if re.match(r'^m[0-9]_commodity_group$', tag.name, flags=re.I):
        return True
    # 第一个模式没匹配到,再检查第二个
    elif re.match(r'^m[0-9]_region_group$', tag.name, flags=re.I):
        return True
    # 前两个都没匹配到,最后检查第三个
    elif re.match(r'^m[0-9]_attribute_group$', tag.name, flags=re.I):
        return True
    # 都不匹配就返回False
    return False

# 先收集所有符合任意模式的标签,再筛选出第一个匹配模式的所有结果
all_matches = soup.find_all(priority_tag_match)
final_results = []
if all_matches:
    first_tag = all_matches[0]
    # 判断第一个匹配项属于哪个模式,筛选同模式的结果
    if re.match(r'^m[0-9]_commodity_group$', first_tag.name, flags=re.I):
        final_results = [t for t in all_matches if re.match(r'^m[0-9]_commodity_group$', t.name, flags=re.I)]
    elif re.match(r'^m[0-9]_region_group$', first_tag.name, flags=re.I):
        final_results = [t for t in all_matches if re.match(r'^m[0-9]_region_group$', t.name, flags=re.I)]
    else:
        final_results = [t for t in all_matches if re.match(r'^m[0-9]_attribute_group$', t.name, flags=re.I)]

print(len(final_results))  # 输出1,符合预期

为什么原来的方法不行?

正则表达式的|是无优先级的或逻辑,只要标签名匹配任意一个模式就会被选中,而find_all会遍历整个文档树,收集所有符合条件的标签,不会因为前面的模式已经匹配到结果就忽略后面的模式。所以你之前的代码会返回所有符合三个模式中任意一个的标签,也就是1+10=11个结果。

内容的提问来源于stack exchange,提问作者Newskooler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:09:05