You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式问题:无法从HTML元素中提取指定属性

搞定特定IFrame的正则匹配与属性提取

嘿,我来帮你实现用正则找到目标iframe并提取它的其他属性,这就给你拆解方案:

一、核心正则表达式

针对你要找的src包含https://www.stackoverflow.com/share/{s}(这里的{s}指代任意后续内容)的iframe,我写了个带捕获组的正则,能精准匹配并提取height和width属性:

<iframe\s+.*?height="(\d+)".*?width="(\d+)".*?src="https://www\.stackoverflow\.com/share.*?"\s*/?>

正则各部分说明:

  • <iframe\s+:匹配iframe标签的开头和后续空格
  • .*?height="(\d+)":非贪婪匹配到height属性,捕获组提取具体数值
  • .*?width="(\d+)":同理,捕获width的数值
  • src="https://www\.stackoverflow\.com/share.*?":精准匹配符合要求的src路径,.*?兼容任意后续内容
  • \s*/?>:匹配标签末尾的空格和闭合符号,兼容自闭合和非自闭合的写法

二、实际代码示例(以Python为例)

下面是直接能用的代码,把你的HTML内容放进去就能提取属性:

import re

# 你的HTML内容
html = '''<iframe height="100" width="200" src="https://www.stackoverflow.com/share"></iframe> <iframe height="100" width="200" src="https://www.google.com/share"></iframe> <iframe height="100" width="200" src="https://www.yahoo.com/share"></iframe>'''

# 正则匹配
pattern = r'<iframe\s+.*?height="(\d+)".*?width="(\d+)".*?src="https://www\.stackoverflow\.com/share.*?"\s*/?>'
results = re.findall(pattern, html)

# 输出提取到的属性
for height, width in results:
    print(f"找到目标IFrame:height={height},width={width}")

运行后会输出:找到目标IFrame:height=100,width=200

三、更靠谱的替代方案(推荐)

正则匹配HTML其实有局限性,比如属性顺序变了、多了其他属性(比如class)就可能失效。如果场景复杂,更推荐用HTML解析库BeautifulSoup,代码更鲁棒:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'html.parser')
# 找到所有src包含目标路径的iframe
target_iframes = soup.find_all('iframe', src=lambda x: x and 'https://www.stackoverflow.com/share' in x)

for iframe in target_iframes:
    print(f"找到目标IFrame:height={iframe.get('height')},width={iframe.get('width')}")

内容的提问来源于stack exchange,提问作者Phorce

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:22:41