Python:如何拆分字符串两个位置并精准提取目标URL
解决方案:提取HTML中的目标URL
不用拆分字符串,直接用正则匹配更精准高效——你的目标是从HTML的href属性里提取被双引号包裹的URL,用正则可以直接定位到目标内容:
import re string = '<p>The brown dog jumped over the... <a href="https://google.com" target="something">... but then splashed in the water<p>' # 匹配以https开头、到双引号结束的内容 match_result = re.search(r'https://[^"]+', string) if match_result: target_url = match_result.group() print(target_url) # 输出:https://google.com
正则说明:
https://:精准匹配目标URL的起始标识[^"]+:匹配除双引号外的任意字符,直到遇到第一个双引号停止,刚好对应href属性里的完整URL
如果字符串里存在多个同类URL,可改用re.findall(r'https://[^"]+', string),会返回所有匹配到的URL组成的列表。
内容的提问来源于stack exchange,提问作者Justin Bertsch
相关产品推荐
相关产品推荐

