You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式从HTML行中提取行星名称?

问题描述

我从HTML文件中得到了如下代码行:

src="https://www.com/seek-images/seek-icons/horoskop-pluto.gif" alt="" />Pluto</div>\n

我尝试用Python正则表达式提取其中的Pluto(或任意行星名称),使用[\w]+作为正则的代码如下:

match = re.search(regexeg, line)
if match:
        print "begin"
        print match.groups()
        for group in match.groups():
                print ('{} '.format(group))

但匹配组始终为空,求解决办法。

解决方案

问题核心有两点:

  1. [\w]+没有定义捕获组,match.groups()只会返回捕获组内容,无括号包裹就不会有结果。
  2. 目标行星名在两处出现(文件名里的小写pluto、标签后的Pluto),需要针对性匹配。

以下是几种实用的提取方案:

提取标签后的行星名称(如示例中的Pluto)

匹配/>*与*</div>之间的单词,正则需包含捕获组:

import re

line = 'src=\"https://www.com/seek-images/seek-icons/horoskop-*pluto.gif*\" alt=\"\" />*Pluto*&lt;/div&gt;\n'
regex = r'/\>\*(\w+)\*\<\/div'
match = re.search(regex, line)
if match:
    print(match.group(1))  # 输出 Pluto

提取文件名中的行星名称(如示例中的pluto)

匹配horoskop-*与.gif*之间的内容:

import re

line = 'src=\"https://www.com/seek-images/seek-icons/horoskop-*pluto.gif*\" alt=\"\" />*Pluto*&lt;/div&gt;\n'
regex = r'horoskop-\*(\w+)\.gif'
match = re.search(regex, line)
if match:
    print(match.group(1))  # 输出 pluto

通用匹配所有可能的行星名称

如果需要同时提取两处的名称,可匹配所有被*包裹的单词(含文件名内的):

import re

line = 'src=\"https://www.com/seek-images/seek-icons/horoskop-*pluto.gif*\" alt=\"\" />*Pluto*&lt;/div&gt;\n'
matches = re.findall(r'\*(\w+)\*(?:\.gif)?', line)
# 去重后输出
for name in set(matches):
    print(name)

重要提示

  • 捕获组必须用()包裹,match.group(1)取第一个捕获组内容,match.groups()返回所有捕获组元组。
  • 正则中的特殊字符(如*、.)需要用\转义,使用原始字符串r''可避免转义混乱。

内容的提问来源于stack exchange,提问作者Class-1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 20:35:06