如何用Python获取动态加载的iframe页面内容?
解决Python爬取动态加载iframe内容的问题
这个问题我之前帮不少人解决过,核心原因是你用的requests只能抓取页面初始的静态HTML,而这个页面的iframe内容是页面加载完成后,通过JavaScript动态发起请求获取并填充进去的,所以BeautifulSoup解析初始页面时,iframe里还是空白的占位内容。给你两个靠谱的解决办法:
方法一:直接请求iframe的真实数据源接口(效率最高)
这是最优解,不需要模拟浏览器,直接抓后台接口数据:
- 打开浏览器开发者工具(按F12),切换到「网络」标签页;
- 刷新目标页面,筛选「XHR」或「Fetch」类型的请求,找和token余额相关的接口(比如URL里带
tokenholders或balances的); - 复制这个接口的URL、请求头(尤其是
User-Agent和Referer,很多网站会校验这些),直接用requests请求就能拿到原始数据。
针对你提供的Etherscan页面,我整理了示例代码(你可以自己去开发者工具里确认最新接口细节):
import requests from bs4 import BeautifulSoup # 模拟浏览器请求头,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://etherscan.io/token/0x168296bb09e24a88805cb9c33356536b980d3fc5#balances' } # 替换成你找到的真实接口URL,p=1代表第一页,多页数据可循环修改p的值 api_url = "https://etherscan.io/token/generic-tokenholders2?m=normal&a=0x168296bb09e24a88805cb9c33356536b980d3fc5&s=999999999999999999999999999&p=1" response = requests.get(api_url, headers=headers) # 解析返回的HTML内容 soup = BeautifulSoup(response.text, 'html.parser') # 提取表格数据 rows = soup.select('table tr') for row in rows[1:]: # 跳过表头行 print(row.get_text(strip=True, separator=' | '))
方法二:用无头浏览器模拟完整页面加载
如果找不到接口或者接口有复杂验证,就用Selenium/Playwright模拟浏览器加载,让iframe自动填充内容:
这里以Selenium为例,先安装依赖:
pip install selenium
然后编写代码:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup # 配置无头浏览器(不弹出可视化窗口) options = Options() options.add_argument('--headless=new') options.add_argument('--disable-gpu') options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') # 启动浏览器 driver = webdriver.Chrome(options=options) driver.get("https://etherscan.io/token/0x168296bb09e24a88805cb9c33356536b980d3fc5#balances") # 等待iframe加载完成(隐式等待10秒,也可以用显式等待更精准) driver.implicitly_wait(10) # 切换到目标iframe(通过xpath定位,也可以用id/name属性) iframe = driver.find_element('xpath', '//iframe[contains(@src, "generic-tokenholders")]') driver.switch_to.frame(iframe) # 获取iframe内的HTML并解析 iframe_html = driver.page_source soup = BeautifulSoup(iframe_html, 'html.parser') # 提取你需要的内容,比如余额表格 table = soup.find('table', class_='table table-md-text-normal') if table: for row in table.find_all('tr')[1:]: cols = row.find_all('td') if cols: rank = cols[0].get_text(strip=True) address = cols[1].get_text(strip=True) balance = cols[2].get_text(strip=True) print(f"排名: {rank} | 地址: {address} | 余额: {balance}") # 关闭浏览器 driver.quit()
注意事项
- 方法一中要注意接口的分页参数(比如
p=1),如果需要爬取多页数据,循环修改分页值即可; - 方法二中Etherscan可能会检测自动化工具,必要时可以添加代理、随机化
User-Agent,或者增加等待时间避免被封。
内容的提问来源于stack exchange,提问作者j.doe
相关产品推荐
相关产品推荐

