You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python获取动态加载的iframe页面内容?

解决Python爬取动态加载iframe内容的问题

这个问题我之前帮不少人解决过,核心原因是你用的requests只能抓取页面初始的静态HTML,而这个页面的iframe内容是页面加载完成后,通过JavaScript动态发起请求获取并填充进去的,所以BeautifulSoup解析初始页面时,iframe里还是空白的占位内容。给你两个靠谱的解决办法:

方法一:直接请求iframe的真实数据源接口(效率最高)

这是最优解,不需要模拟浏览器,直接抓后台接口数据:

  • 打开浏览器开发者工具(按F12),切换到「网络」标签页;
  • 刷新目标页面,筛选「XHR」或「Fetch」类型的请求,找和token余额相关的接口(比如URL里带tokenholders或balances的);
  • 复制这个接口的URL、请求头(尤其是User-Agent和Referer,很多网站会校验这些),直接用requests请求就能拿到原始数据。

针对你提供的Etherscan页面,我整理了示例代码(你可以自己去开发者工具里确认最新接口细节):

import requests
from bs4 import BeautifulSoup

# 模拟浏览器请求头,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Referer': 'https://etherscan.io/token/0x168296bb09e24a88805cb9c33356536b980d3fc5#balances'
}

# 替换成你找到的真实接口URL,p=1代表第一页,多页数据可循环修改p的值
api_url = "https://etherscan.io/token/generic-tokenholders2?m=normal&a=0x168296bb09e24a88805cb9c33356536b980d3fc5&s=999999999999999999999999999&p=1"
response = requests.get(api_url, headers=headers)

# 解析返回的HTML内容
soup = BeautifulSoup(response.text, 'html.parser')
# 提取表格数据
rows = soup.select('table tr')
for row in rows[1:]:  # 跳过表头行
    print(row.get_text(strip=True, separator=' | '))

方法二:用无头浏览器模拟完整页面加载

如果找不到接口或者接口有复杂验证,就用Selenium/Playwright模拟浏览器加载,让iframe自动填充内容:
这里以Selenium为例,先安装依赖:

pip install selenium

然后编写代码:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

# 配置无头浏览器(不弹出可视化窗口)
options = Options()
options.add_argument('--headless=new')
options.add_argument('--disable-gpu')
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

# 启动浏览器
driver = webdriver.Chrome(options=options)
driver.get("https://etherscan.io/token/0x168296bb09e24a88805cb9c33356536b980d3fc5#balances")

# 等待iframe加载完成(隐式等待10秒,也可以用显式等待更精准)
driver.implicitly_wait(10)

# 切换到目标iframe(通过xpath定位,也可以用id/name属性)
iframe = driver.find_element('xpath', '//iframe[contains(@src, "generic-tokenholders")]')
driver.switch_to.frame(iframe)

# 获取iframe内的HTML并解析
iframe_html = driver.page_source
soup = BeautifulSoup(iframe_html, 'html.parser')

# 提取你需要的内容,比如余额表格
table = soup.find('table', class_='table table-md-text-normal')
if table:
    for row in table.find_all('tr')[1:]:
        cols = row.find_all('td')
        if cols:
            rank = cols[0].get_text(strip=True)
            address = cols[1].get_text(strip=True)
            balance = cols[2].get_text(strip=True)
            print(f"排名: {rank} | 地址: {address} | 余额: {balance}")

# 关闭浏览器
driver.quit()

注意事项

  • 方法一中要注意接口的分页参数(比如p=1),如果需要爬取多页数据,循环修改分页值即可;
  • 方法二中Etherscan可能会检测自动化工具,必要时可以添加代理、随机化User-Agent,或者增加等待时间避免被封。

内容的提问来源于stack exchange,提问作者j.doe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:33:37