如何用Python正则表达式从HTML行中提取行星名称?
问题描述
我从HTML文件中得到了如下代码行:
src="https://www.com/seek-images/seek-icons/horoskop-pluto.gif" alt="" />Pluto</div>\n
我尝试用Python正则表达式提取其中的Pluto(或任意行星名称),使用[\w]+作为正则的代码如下:
match = re.search(regexeg, line) if match: print "begin" print match.groups() for group in match.groups(): print ('{} '.format(group))
但匹配组始终为空,求解决办法。
解决方案
问题核心有两点:
[\w]+没有定义捕获组,match.groups()只会返回捕获组内容,无括号包裹就不会有结果。- 目标行星名在两处出现(文件名里的小写
pluto、标签后的Pluto),需要针对性匹配。
以下是几种实用的提取方案:
提取标签后的行星名称(如示例中的Pluto)
匹配/>*与*</div>之间的单词,正则需包含捕获组:
import re line = 'src=\"https://www.com/seek-images/seek-icons/horoskop-*pluto.gif*\" alt=\"\" />*Pluto*</div>\n' regex = r'/\>\*(\w+)\*\<\/div' match = re.search(regex, line) if match: print(match.group(1)) # 输出 Pluto
提取文件名中的行星名称(如示例中的pluto)
匹配horoskop-*与.gif*之间的内容:
import re line = 'src=\"https://www.com/seek-images/seek-icons/horoskop-*pluto.gif*\" alt=\"\" />*Pluto*</div>\n' regex = r'horoskop-\*(\w+)\.gif' match = re.search(regex, line) if match: print(match.group(1)) # 输出 pluto
通用匹配所有可能的行星名称
如果需要同时提取两处的名称,可匹配所有被*包裹的单词(含文件名内的):
import re line = 'src=\"https://www.com/seek-images/seek-icons/horoskop-*pluto.gif*\" alt=\"\" />*Pluto*</div>\n' matches = re.findall(r'\*(\w+)\*(?:\.gif)?', line) # 去重后输出 for name in set(matches): print(name)
重要提示
- 捕获组必须用
()包裹,match.group(1)取第一个捕获组内容,match.groups()返回所有捕获组元组。 - 正则中的特殊字符(如
*、.)需要用\转义,使用原始字符串r''可避免转义混乱。
内容的提问来源于stack exchange,提问作者Class-1
相关产品推荐
相关产品推荐

