关于从Bscscan爬取代币顶级持有者地址及占比的技术求助(含多轮尝试代码)
Hey there, let's break down how to fix your Bscscan scraping attempts and get that token holder data (address 0x7754c0584372D29510C019136220f91e25a8f706) you need. Bscscan has strict anti-scraping measures, so your raw requests are likely getting blocked, missing dynamic content, or using invalid selectors. Here's how to tweak your code and some more reliable approaches:
First: Basic Anti-Scraping Mitigations
Before diving into fixes, add these to all your requests to mimic a real browser:
- Use a persistent session to maintain cookies
- Fill out complete request headers (not just User-Agent)
- Add random delays between requests to avoid rate-limiting
Fixes for Your Four Attempts
Attempt 1: Raw Requests + XPath
Problems: No anti-scraping headers, Bscscan might block you; XPath targets tbody which is often dynamically injected by JS (so it won't exist in the raw HTML).
Improved Code:
import requests from bs4 import BeautifulSoup from lxml import etree import time token = "0x7754c0584372D29510C019136220f91e25a8f706" url = f"https://bscscan.com/token/{token}#balances" # Use a session to persist cookies session = requests.Session() headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "https://bscscan.com/", "Accept-Language": "en-US,en;q=0.9" } # Add delay to avoid rate limits time.sleep(2) response = session.get(url, headers=headers) if response.status_code != 200: print(f"Request failed with status code: {response.status_code}") print("You might need to handle a CAPTCHA or IP block") else: soup = BeautifulSoup(response.content, "html.parser") d = etree.HTML(str(soup)) # Skip tbody (dynamic) and target tr directly percentage = d.xpath('//*[@id="maintable"]/div[3]/table/tr[1]/td[4]/text()') address_href = d.xpath('//*[@id="maintable"]/div[3]/table/tr[1]/td[2]/span/a/@href') print("Top holder percentage:", percentage[0].strip() if percentage else "Not found") print("Top holder address:", address_href[0].split('/')[-1] if address_href else "Not found")
Attempt 2: Direct etree.parse(url)
Problems: etree.parse uses a raw HTTP request without anti-scraping headers, so Bscscan will block it. It also can't handle dynamic content.
Fix: Use the session-based request from Attempt 1, then parse the response content instead of the URL directly.
Attempt 3: Invalid BeautifulSoup Selectors
Problems: soup.find_all('<td>1</td>') is invalid syntax (find_all takes tag names/attributes, not raw HTML); _parent isn't a valid attribute to target.
Improved Code (build on the session from Attempt 1):
# After getting the soup object rows = soup.select('#maintable div.table-responsive table tr') if len(rows) > 1: # Skip the header row first_row = rows[1] address_elem = first_row.select_one('td:nth-child(2) span a') percentage_elem = first_row.select_one('td:nth-child(4)') if address_elem and percentage_elem: print("Top holder address:", address_elem['href'].split('/')[-1]) print("Top holder percentage:", percentage_elem.get_text(strip=True))
Attempt 4: Hardcoded sid Parameter
Problems: The sid in your URL is dynamically generated per session—using a hardcoded value will fail. Your CSS selectors are also incorrect (td[3] should be td:nth-child(3), and _parent isn't a valid attribute).
Improved Code:
import requests from parsel import Selector import time import urllib.parse token = "0x7754c0584372D29510C019136220f91e25a8f706" main_url = f"https://bscscan.com/token/{token}#balances" session = requests.Session() headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "https://bscscan.com/", "Accept-Language": "en-US,en;q=0.9" } # First, get the dynamic sid from the main token page time.sleep(2) main_response = session.get(main_url, headers=headers) if main_response.status_code != 200: print("Failed to load main token page") else: sel_main = Selector(main_response.text) holder_link = sel_main.css('a[href*="generic-tokenholders2"]::attr(href)').get() if holder_link: # Extract the sid parameter from the link query_params = urllib.parse.parse_qs(holder_link.split('?')[1]) sid = query_params.get('sid', [None])[0] if sid: # Build the valid holder list URL holder_url = f"https://bscscan.com/token/generic-tokenholders2?m=normal&a={token}&s=100000000000000000000000000&sid={sid}&p=1" time.sleep(2) holder_response = session.get(holder_url, headers=headers) sel_holder = Selector(holder_response.text) # Correct CSS selectors addresses = sel_holder.css('tr td:nth-child(2) span a::attr(href)').extract() cleaned_addresses = [addr.split('/')[-1] for addr in addresses] percentages = sel_holder.css('tr td:nth-child(4)::text').extract() cleaned_percentages = [p.strip() for p in percentages if p.strip()] print("Top 5 holder addresses:", cleaned_addresses[:5]) print("Top 5 holder percentages:", cleaned_percentages[:5]) else: print("Couldn't extract dynamic sid parameter") else: print("Couldn't find holder list link on main page")
Easier Alternative: Headless Browsers
For more reliable scraping of dynamic content (and to avoid dealing with raw requests), use a headless browser like Playwright or Selenium. It mimics real user behavior, which bypasses most basic anti-scraping measures.
Playwright Example:
from playwright.sync_api import sync_playwright token = "0x7754c0584372D29510C019136220f91e25a8f706" url = f"https://bscscan.com/token/{token}#balances" with sync_playwright() as p: # Launch browser with anti-detection flags browser = p.chromium.launch( headless=True, args=["--disable-blink-features=AutomationControlled"] ) context = browser.new_context( user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" ) page = context.new_page() # Wait for network to be idle to ensure content loads page.goto(url, wait_until="networkidle") # Wait for the holder table to load page.wait_for_selector('#maintable div.table-responsive table tr') # Extract first row data first_row = page.locator('#maintable div.table-responsive table tr').nth(1) address_href = first_row.locator('td:nth-child(2) span a').get_attribute('href') address = address_href.split('/')[-1] percentage = first_row.locator('td:nth-child(4)').text_content().strip() print(f"Top Holder Address: {address}") print(f"Top Holder Percentage: {percentage}") browser.close()
Note: Even with headless browsers, you'll need to control request frequency to avoid IP blocks. If you hit CAPTCHAs, you might need manual intervention or a CAPTCHA-solving service (though that adds complexity).
内容的提问来源于stack exchange,提问作者BJonas88

