如何在BeautifulSoup的find_all方法中使用正则表达式实现带优先级的匹配
实现带优先级的标签匹配(按顺序匹配,匹配到即停止后续模式)
你遇到的问题是因为正则表达式的|是逻辑或,只要标签名符合任意一个模式就会被选中,而BeautifulSoup的find_all会遍历所有标签并收集所有匹配项,不会自动帮你实现“优先匹配第一个模式,成功就忽略后面”的逻辑。
下面给你两种可行的解决方案,都能实现你要的优先级匹配效果:
方案一:分步查找(最直观易维护)
按优先级顺序依次查找,找到第一个有结果的模式就停止,直接返回该结果:
from bs4 import BeautifulSoup import re xml_doc = """ <m3_commodity_group commodity3="Oilseeds"><m3_year_group_Collection><m3_year_group market_year3="2011/12"><m3_month_group_Collection><m3_month_group forecast_month3=""><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell cell_value3="353.93"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell cell_value3="429.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell cell_value3="73.59"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell cell_value3="345.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell cell_value3="59.03"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection><m3_value_group_Collection><m3_value_group><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection></m3_value_group></m3_value_group_Collection></m3_month_group></m3_month_group_Collection></m3_year_group></m3_year_group_Collection></m3_commodity_group> """ soup = BeautifulSoup(xml_doc, "xml") # 定义优先级模式列表,顺序就是优先级 priority_patterns = [ r'^m[0-9]_commodity_group$', r'^m[0-9]_region_group$', r'^m[0-9]_attribute_group$' ] matched_results = [] for pattern in priority_patterns: current_matches = soup.find_all(re.compile(pattern, flags=re.I)) if current_matches: matched_results = current_matches break # 找到第一个匹配的模式,直接停止后续查找 print(len(matched_results)) # 输出1,和你预期的一致
方案二:自定义匹配函数(适合一次性传递逻辑)
利用BeautifulSoup支持传入自定义函数作为匹配条件的特性,在函数内部实现优先级判断:
from bs4 import BeautifulSoup import re xml_doc = """ <m3_commodity_group commodity3="Oilseeds"><m3_year_group_Collection><m3_year_group market_year3="2011/12"><m3_month_group_Collection><m3_month_group forecast_month3=""><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell cell_value3="353.93"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell cell_value3="429.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell cell_value3="73.59"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell cell_value3="345.49"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell cell_value3="59.03"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection><m3_value_group_Collection><m3_value_group><m3_attribute_group_Collection><m3_attribute_group attribute3="Output"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Supply"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Trade"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Total Use 2/"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group><m3_attribute_group attribute3="Ending Stocks"><Textbox40><Cell Textbox44="filler"/></Textbox40></m3_attribute_group></m3_attribute_group_Collection></m3_value_group></m3_value_group_Collection></m3_month_group></m3_month_group_Collection></m3_year_group></m3_year_group_Collection></m3_commodity_group> """ soup = BeautifulSoup(xml_doc, "xml") def priority_tag_match(tag): # 按优先级检查:先看是否匹配第一个模式 if re.match(r'^m[0-9]_commodity_group$', tag.name, flags=re.I): return True # 第一个模式没匹配到,再检查第二个 elif re.match(r'^m[0-9]_region_group$', tag.name, flags=re.I): return True # 前两个都没匹配到,最后检查第三个 elif re.match(r'^m[0-9]_attribute_group$', tag.name, flags=re.I): return True # 都不匹配就返回False return False # 先收集所有符合任意模式的标签,再筛选出第一个匹配模式的所有结果 all_matches = soup.find_all(priority_tag_match) final_results = [] if all_matches: first_tag = all_matches[0] # 判断第一个匹配项属于哪个模式,筛选同模式的结果 if re.match(r'^m[0-9]_commodity_group$', first_tag.name, flags=re.I): final_results = [t for t in all_matches if re.match(r'^m[0-9]_commodity_group$', t.name, flags=re.I)] elif re.match(r'^m[0-9]_region_group$', first_tag.name, flags=re.I): final_results = [t for t in all_matches if re.match(r'^m[0-9]_region_group$', t.name, flags=re.I)] else: final_results = [t for t in all_matches if re.match(r'^m[0-9]_attribute_group$', t.name, flags=re.I)] print(len(final_results)) # 输出1,符合预期
为什么原来的方法不行?
正则表达式的|是无优先级的或逻辑,只要标签名匹配任意一个模式就会被选中,而find_all会遍历整个文档树,收集所有符合条件的标签,不会因为前面的模式已经匹配到结果就忽略后面的模式。所以你之前的代码会返回所有符合三个模式中任意一个的标签,也就是1+10=11个结果。
内容的提问来源于stack exchange,提问作者Newskooler
相关产品推荐
相关产品推荐

