如何使用Python从本地HTML文件中提取加密货币补充数据到输出结果
实现思路
你需要的所有代币信息都已经存在本地HTML文件中,无需额外请求第三方站点,直接按HTML结构提取即可:
- 每个代币的完整信息都封装在独立的
class="text"的div节点内,先遍历所有这类节点,而不是直接查找a标签,能避免数据错位 - 对每个节点,先提取BscScan链接,再分别匹配对应的名称、总供应量、流动性、持币地址数、转账数字段
- 总供应量的代币符号可以直接提取相邻的strong标签内容拼接即可
完整实现代码
from bs4 import BeautifulSoup import re srcfile = 'sourcefile.html' # 用with open管理文件句柄,避免资源泄漏 with open(srcfile, 'r', encoding="utf-8") as f: soup = BeautifulSoup(f, "html.parser") # 遍历每个代币对应的独立信息区块 for token_block in soup.find_all("div", class_="text"): # 提取BscScan代币链接 bsc_link = token_block.find("a", href=re.compile('https://bscscan\.com/token/'))["href"] # 提取代币名称 name = token_block.find("strong", string=re.compile("Name")).next_sibling.strip().lstrip(": ") # 提取总供应量+代币符号 supply_ele = token_block.find("strong", string=re.compile("Total Supply")) supply_val = supply_ele.next_sibling.strip().lstrip(": ") supply_symbol = supply_ele.find_next_sibling("strong").text.strip() total_supply = f"{supply_val} {supply_symbol}" # 提取流动性数据 liquidity = token_block.find("strong", string=re.compile("Liquidity")).next_sibling.strip().lstrip(": ") # 提取持币地址数 holders = token_block.find("strong", string=re.compile("Holders")).next_sibling.strip().lstrip(": ") # 提取转账数 transfers = token_block.find("strong", string=re.compile("Transfers")).next_sibling.strip().lstrip(": ") # 格式化输出,对齐效果和你要求的一致 print(f"Name: {name:<35} Total Supply: {total_supply}") print(f"Liquidity: {liquidity:<30} Holders: {holders:<10} Transfers: {transfers}") print(f" BscScan: {bsc_link}\n")
运行效果说明
代码输出完全匹配你期望的格式,所有数据均从本地HTML直接提取,运行速度快也不会有反爬限制。
内容的提问来源于stack exchange,提问作者rbutrnz
相关产品推荐
相关产品推荐

