Python如何提取td标签内的指定链接 现有代码输出不符求调整
代码修正方案
存在的问题
- 目标元素查找范围错误:原代码使用全局页面soup
soupblockdetails查找a标签,只会返回整个页面第一个匹配的元素,不是当前行对应交易的目标地址,因此输出了错误的Validator: Stake2me内容 - 输出格式不匹配:原打印语句未加入地址变量,无法输出预期的三段结构内容
修正后代码
from bs4 import BeautifulSoup from time import sleep import requests headers = {"User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:92.0) Gecko/20100101 Firefox/92.0"} urllink = "https://bscscan.com/txs?block=11711353&ps=100&p=1" reqblockdetails = requests.get(urllink, headers=headers, timeout=5) soupblockdetails = BeautifulSoup(reqblockdetails.content, 'html.parser') rowsblockdetails = soupblockdetails.findAll('table')[0].findAll('tr') sleep(1) for row in rowsblockdetails[1:]: txnhash = row.find_all('td')[1].text[0:] txnhashdetails = txnhash.strip() destination_td = row.find_all('td')[8] destination = destination_td.text.strip() if destination == "CoinOne: CONE Token": # 从当前行的目标td中找a标签,拆分href获取地址 dest_a = destination_td.find('a', attrs={'class': 'hash-tag text-truncate'}) urldest = dest_a['href'].split('/')[-1] print (" {:>1} {:<5} {}".format(txnhashdetails, destination, urldest))
核心调整点
- 把查找a标签的范围从全局页面soup改为当前遍历行的对应td,保证取到的是当前交易的匹配元素
- 从a标签的href属性中拆分出末尾的0x开头合约地址,符合预期输出要求
- 调整打印语句的格式化参数,补全地址字段的输出位置
内容的提问来源于stack exchange,提问作者rbutrnz
相关产品推荐
相关产品推荐

