用Python BeautifulSoup限制自动刷新页面抓取,去重并保障区块连续性
解决方案
核心实现逻辑:
- 新增
last_processed_block变量记录已经处理完成的最高区块号,初始值设为0 - 每次拉取页面后遍历交易行,仅处理区块号大于
last_processed_block的交易,一旦碰到小于等于已处理最高区块号的交易就终止当前页的遍历,避免冗余计算 - 每次处理完当前页的新交易后,更新
last_processed_block为本次扫描到的最高区块号,确保不会漏块也不会重复处理
修改后的完整代码:
from bs4 import BeautifulSoup from time import sleep import re, requests trim = re.compile(r'[^\d,.]+') url = "https://bscscan.com/txs?a=0x10ed43c718714eb63d5aa57b78b54704e256024e&ps=100&p=1" baseurl = 'https://bscscan.com/tx/' header = {"User-Agent": "Mozilla/5.0"} scans = 0 # 记录已处理的最高区块 last_processed_block = 0 while True: scans += 1 reqtxsInternal = requests.get(url, headers=header, timeout=2) souptxsInternal = BeautifulSoup(reqtxsInternal.content, 'html.parser') blocktxsInternal = souptxsInternal.findAll('table')[0].findAll('tr') # 记录本次扫描到的最高区块 current_max_block = last_processed_block for row in blocktxsInternal[1:]: txnhash = row.find_all('td')[1].text[0:] txnhashdetails = txnhash.strip() block = int(row.find_all('td')[3].text.strip()) # 碰到已处理过的区块直接终止遍历,后面都是更早的区块 if block <= last_processed_block: break # 更新本次扫描的最高区块号 if block > current_max_block: current_max_block = block value = row.find_all('td')[9].text[0:] amount = trim.sub('', value).replace(",", "") transval = float(amount) if transval >= 1: print(f"Doing something with the data -> {block} {transval}") # 更新全局已处理最高区块号 last_processed_block = current_max_block print (" -> Whole Page Scanned: ", scans) sleep(1)
运行后完全符合预期效果:
- 第一次扫描仅处理最高区块10186993的所有符合条件的交易,碰到10186992就停止遍历
- 第二次扫描检测到新的最高区块10186994,处理完该区块的所有符合条件交易后碰到10186993就停止,不会输出重复内容
- 如果出现区块跳跃(比如上次处理到10186993,下次直接出现10186995),也会自动处理10186994和10186995两个块的交易,保证区块连续性无遗漏
内容的提问来源于stack exchange,提问作者rbutrnz
相关产品推荐
相关产品推荐

