You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则无法正确识别WPRM食谱字符串分隔符的修复方案咨询

问题

我正在编写一款可自动缩放WPRM(WordPress Recipe Maker)食谱的程序,计划通过正则表达式从食谱打印版HTML代码中提取配料的数量、单位、名称信息。相关信息被类似...ingredient-attribute">和</span>的标签包裹。

我编写了如下代码:

ingredients = re.split(r'<li', recipestr_cut) 
    
for x in ingredients:
    
    out = [] # initialize a single element of the recipe as a list
    
    amount_m = re.search(r'ingredient-amount"\>(.+)\<\/span>', x)
    if amount_m:
        out.append(amount_m.group(1))
    unit_m = re.search(r'ingredient-unit"\>(.+)\<\/span>', x)
    if unit_m:
        out.append(unit_m.group(1))
    ingredient_m = re.search(r'ingredient-name"\>(.+)\<\/span>', x)
    if ingredient_m:
        out.append(ingredient_m.group(1))
    
    if len(out) > 0:
        recipe_readable.append(out)

该代码在简单食谱(如曲奇食谱)上可正常提取,返回结果如下:

[['3', 'cups', '(380 grams) all-purpose flour'], ['1', 'teaspoon', 'baking soda'], ['1', 'teaspoon', 'fine sea salt'], ['2', 'sticks (227 grams) unsalted butter, at cool room temperature (67°F)'], ['1/2', 'cup', '(100 grams) granulated sugar'], ['1 1/4', 'cups', '(247 grams) lightly packed light brown sugar'], ['2', 'teaspoons', 'vanilla'], ['2', 'large eggs, at room temperature'], ['2', 'cups', '(340 grams) semisweet chocolate chips']]

但在处理复杂食谱(如猪肉包食谱)时,正则无法正确识别</span>分隔符,提取结果包含多余标签内容,例如:

[['2/3</span> <span class="wprm-recipe-ingredient-unit">cup</span> <span class="wprm-recipe-ingredient-name">heavy cream', 'cup</span> <span class="wprm-recipe-ingredient-name">heavy cream</span> <span class="wprm-recipe-ingredient-notes wprm-recipe-ingredient-notes-faded">(at room temperature)', ...

预期结果应为[['2/3','cup','heavy cream'],...]。我刚接触正则表达式,请问该如何修复这个问题?

解决方法

1. 修复正则的贪婪匹配问题

你的正则中使用的.+是贪婪匹配,会尽可能抓取最长的内容。在复杂HTML结构里,它会从第一个匹配的标签起始处,一直匹配到最后一个</span>,导致把后续的标签内容也包含进来。

只需将.+改为.+?(非贪婪匹配),让正则匹配到第一个</span>就停止:

修改后的正则代码:

amount_m = re.search(r'ingredient-amount"\>(.+?)\<\/span>', x)
unit_m = re.search(r'ingredient-unit"\>(.+?)\<\/span>', x)
ingredient_m = re.search(r'ingredient-name"\>(.+?)\<\/span>', x)

这样就能精准提取每个标签内的目标内容,不会混入后续的HTML代码。

2. 更可靠的替代方案:使用HTML解析库

正则处理HTML天生存在局限性,遇到嵌套标签、特殊字符转义时很容易出错。推荐使用BeautifulSoup这类专业的HTML解析库,代码更健壮,维护成本更低。

示例代码如下:

from bs4 import BeautifulSoup

# 先将转义的HTML字符还原为原生标签
soup = BeautifulSoup(
    recipestr_cut.replace('&lt;', '<').replace('&gt;', '>').replace('&quot;', '"'),
    'html.parser'
)

recipe_readable = []
# 遍历所有配料对应的li元素
for li in soup.find_all('li'):
    out = []
    # 通过类名定位并提取内容
    amount = li.find(class_='wprm-recipe-ingredient-amount')
    if amount:
        out.append(amount.get_text(strip=True))
    unit = li.find(class_='wprm-recipe-ingredient-unit')
    if unit:
        out.append(unit.get_text(strip=True))
    name = li.find(class_='wprm-recipe-ingredient-name')
    if name:
        out.append(name.get_text(strip=True))
    if out:
        recipe_readable.append(out)

这个方法无需编写复杂正则,直接通过类名定位元素,能处理各种复杂HTML结构,还会自动处理字符转义问题。

内容的提问来源于stack exchange,提问作者paulina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 08:57:24