如何使用Python和BeautifulSoup抓取网页中iframe的指定范围数据
问题原因
- 你遇到的
Invalid URL报错是因为未携带浏览器标识请求头直接访问bscscan,站点反爬机制返回的是拦截页面,无法提取到有效iframe的src属性,拼接后url非法触发报错。 - 现有代码还存在未定义变量问题:遍历
rowsblockdetails前,没有从iframe返回的页面内容中提取对应表格行变量,后续运行也会报错。 - bscscan对iframe请求也有校验,需要携带对应请求头才能拿到正常的持仓数据。
修正后可运行代码
import requests from bs4 import BeautifulSoup # 配置请求头,伪装成浏览器访问 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Referer": "https://bscscan.com/" } s = requests.Session() # 请求主页面携带请求头 main_url = "https://bscscan.com/token/0xe56842ed550ff2794f010738554db45e60730371#balances" r = s.get(main_url, headers=headers) soup = BeautifulSoup(r.content, "html.parser") iframe_src = soup.select_one("#tokeholdersiframe").attrs["src"] # 请求iframe页面也携带请求头 r = s.get(f"https://bscscan.com{iframe_src}", headers=headers) soup = BeautifulSoup(r.content, "html.parser") # 提取表格行,解决未定义变量问题 rowsblockdetails = soup.select("table.table tr") for row in rowsblockdetails[1:]: tds = row.find_all('td') if len(tds) < 4: continue rank = tds[0].text.strip() address = tds[1].text.strip() amount = tds[2].text.strip() percentage = tds[3].text.strip() # 提取合约标识 tag = tds[1].select_one(".text-secondary") tag_text = tag.text.strip() if tag else "" print(f" {rank:<3} {address:<45} {amount:<30} {percentage:<10} {tag_text}")
运行说明
代码运行后会直接输出你期望的持仓列表,包含排名、地址、持仓数量、持仓占比、合约标识字段。如果需要调整输出对齐效果,修改print语句中的格式化参数即可。
内容的提问来源于stack exchange,提问作者rbutrnz
相关产品推荐
相关产品推荐

