正则表达式问题:无法从HTML元素中提取指定属性
搞定特定IFrame的正则匹配与属性提取
嘿,我来帮你实现用正则找到目标iframe并提取它的其他属性,这就给你拆解方案:
一、核心正则表达式
针对你要找的src包含https://www.stackoverflow.com/share/{s}(这里的{s}指代任意后续内容)的iframe,我写了个带捕获组的正则,能精准匹配并提取height和width属性:
<iframe\s+.*?height="(\d+)".*?width="(\d+)".*?src="https://www\.stackoverflow\.com/share.*?"\s*/?>
正则各部分说明:
<iframe\s+:匹配iframe标签的开头和后续空格.*?height="(\d+)":非贪婪匹配到height属性,捕获组提取具体数值.*?width="(\d+)":同理,捕获width的数值src="https://www\.stackoverflow\.com/share.*?":精准匹配符合要求的src路径,.*?兼容任意后续内容\s*/?>:匹配标签末尾的空格和闭合符号,兼容自闭合和非自闭合的写法
二、实际代码示例(以Python为例)
下面是直接能用的代码,把你的HTML内容放进去就能提取属性:
import re # 你的HTML内容 html = '''<iframe height="100" width="200" src="https://www.stackoverflow.com/share"></iframe> <iframe height="100" width="200" src="https://www.google.com/share"></iframe> <iframe height="100" width="200" src="https://www.yahoo.com/share"></iframe>''' # 正则匹配 pattern = r'<iframe\s+.*?height="(\d+)".*?width="(\d+)".*?src="https://www\.stackoverflow\.com/share.*?"\s*/?>' results = re.findall(pattern, html) # 输出提取到的属性 for height, width in results: print(f"找到目标IFrame:height={height},width={width}")
运行后会输出:找到目标IFrame:height=100,width=200
三、更靠谱的替代方案(推荐)
正则匹配HTML其实有局限性,比如属性顺序变了、多了其他属性(比如class)就可能失效。如果场景复杂,更推荐用HTML解析库BeautifulSoup,代码更鲁棒:
from bs4 import BeautifulSoup soup = BeautifulSoup(html, 'html.parser') # 找到所有src包含目标路径的iframe target_iframes = soup.find_all('iframe', src=lambda x: x and 'https://www.stackoverflow.com/share' in x) for iframe in target_iframes: print(f"找到目标IFrame:height={iframe.get('height')},width={iframe.get('width')}")
内容的提问来源于stack exchange,提问作者Phorce
相关产品推荐
相关产品推荐

