You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于从Bscscan爬取代币顶级持有者地址及占比的技术求助(含多轮尝试代码)

Fixing Bscscan Token Holder Scraping Issues

Hey there, let's break down how to fix your Bscscan scraping attempts and get that token holder data (address 0x7754c0584372D29510C019136220f91e25a8f706) you need. Bscscan has strict anti-scraping measures, so your raw requests are likely getting blocked, missing dynamic content, or using invalid selectors. Here's how to tweak your code and some more reliable approaches:

First: Basic Anti-Scraping Mitigations

Before diving into fixes, add these to all your requests to mimic a real browser:

  • Use a persistent session to maintain cookies
  • Fill out complete request headers (not just User-Agent)
  • Add random delays between requests to avoid rate-limiting

Fixes for Your Four Attempts

Attempt 1: Raw Requests + XPath

Problems: No anti-scraping headers, Bscscan might block you; XPath targets tbody which is often dynamically injected by JS (so it won't exist in the raw HTML).

Improved Code:

import requests
from bs4 import BeautifulSoup
from lxml import etree
import time

token = "0x7754c0584372D29510C019136220f91e25a8f706"
url = f"https://bscscan.com/token/{token}#balances"

# Use a session to persist cookies
session = requests.Session()
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Referer": "https://bscscan.com/",
    "Accept-Language": "en-US,en;q=0.9"
}

# Add delay to avoid rate limits
time.sleep(2)
response = session.get(url, headers=headers)

if response.status_code != 200:
    print(f"Request failed with status code: {response.status_code}")
    print("You might need to handle a CAPTCHA or IP block")
else:
    soup = BeautifulSoup(response.content, "html.parser")
    d = etree.HTML(str(soup))
    # Skip tbody (dynamic) and target tr directly
    percentage = d.xpath('//*[@id="maintable"]/div[3]/table/tr[1]/td[4]/text()')
    address_href = d.xpath('//*[@id="maintable"]/div[3]/table/tr[1]/td[2]/span/a/@href')
    
    print("Top holder percentage:", percentage[0].strip() if percentage else "Not found")
    print("Top holder address:", address_href[0].split('/')[-1] if address_href else "Not found")

Attempt 2: Direct etree.parse(url)

Problems: etree.parse uses a raw HTTP request without anti-scraping headers, so Bscscan will block it. It also can't handle dynamic content.

Fix: Use the session-based request from Attempt 1, then parse the response content instead of the URL directly.

Attempt 3: Invalid BeautifulSoup Selectors

Problems: soup.find_all('<td>1</td>') is invalid syntax (find_all takes tag names/attributes, not raw HTML); _parent isn't a valid attribute to target.

Improved Code (build on the session from Attempt 1):

# After getting the soup object
rows = soup.select('#maintable div.table-responsive table tr')
if len(rows) > 1:  # Skip the header row
    first_row = rows[1]
    address_elem = first_row.select_one('td:nth-child(2) span a')
    percentage_elem = first_row.select_one('td:nth-child(4)')
    
    if address_elem and percentage_elem:
        print("Top holder address:", address_elem['href'].split('/')[-1])
        print("Top holder percentage:", percentage_elem.get_text(strip=True))

Attempt 4: Hardcoded sid Parameter

Problems: The sid in your URL is dynamically generated per session—using a hardcoded value will fail. Your CSS selectors are also incorrect (td[3] should be td:nth-child(3), and _parent isn't a valid attribute).

Improved Code:

import requests
from parsel import Selector
import time
import urllib.parse

token = "0x7754c0584372D29510C019136220f91e25a8f706"
main_url = f"https://bscscan.com/token/{token}#balances"
session = requests.Session()
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Referer": "https://bscscan.com/",
    "Accept-Language": "en-US,en;q=0.9"
}

# First, get the dynamic sid from the main token page
time.sleep(2)
main_response = session.get(main_url, headers=headers)
if main_response.status_code != 200:
    print("Failed to load main token page")
else:
    sel_main = Selector(main_response.text)
    holder_link = sel_main.css('a[href*="generic-tokenholders2"]::attr(href)').get()
    
    if holder_link:
        # Extract the sid parameter from the link
        query_params = urllib.parse.parse_qs(holder_link.split('?')[1])
        sid = query_params.get('sid', [None])[0]
        
        if sid:
            # Build the valid holder list URL
            holder_url = f"https://bscscan.com/token/generic-tokenholders2?m=normal&a={token}&s=100000000000000000000000000&sid={sid}&p=1"
            time.sleep(2)
            holder_response = session.get(holder_url, headers=headers)
            sel_holder = Selector(holder_response.text)
            
            # Correct CSS selectors
            addresses = sel_holder.css('tr td:nth-child(2) span a::attr(href)').extract()
            cleaned_addresses = [addr.split('/')[-1] for addr in addresses]
            percentages = sel_holder.css('tr td:nth-child(4)::text').extract()
            cleaned_percentages = [p.strip() for p in percentages if p.strip()]
            
            print("Top 5 holder addresses:", cleaned_addresses[:5])
            print("Top 5 holder percentages:", cleaned_percentages[:5])
        else:
            print("Couldn't extract dynamic sid parameter")
    else:
        print("Couldn't find holder list link on main page")

Easier Alternative: Headless Browsers

For more reliable scraping of dynamic content (and to avoid dealing with raw requests), use a headless browser like Playwright or Selenium. It mimics real user behavior, which bypasses most basic anti-scraping measures.

Playwright Example:

from playwright.sync_api import sync_playwright

token = "0x7754c0584372D29510C019136220f91e25a8f706"
url = f"https://bscscan.com/token/{token}#balances"

with sync_playwright() as p:
    # Launch browser with anti-detection flags
    browser = p.chromium.launch(
        headless=True,
        args=["--disable-blink-features=AutomationControlled"]
    )
    context = browser.new_context(
        user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    )
    page = context.new_page()
    
    # Wait for network to be idle to ensure content loads
    page.goto(url, wait_until="networkidle")
    # Wait for the holder table to load
    page.wait_for_selector('#maintable div.table-responsive table tr')
    
    # Extract first row data
    first_row = page.locator('#maintable div.table-responsive table tr').nth(1)
    address_href = first_row.locator('td:nth-child(2) span a').get_attribute('href')
    address = address_href.split('/')[-1]
    percentage = first_row.locator('td:nth-child(4)').text_content().strip()
    
    print(f"Top Holder Address: {address}")
    print(f"Top Holder Percentage: {percentage}")
    
    browser.close()

Note: Even with headless browsers, you'll need to control request frequency to avoid IP blocks. If you hit CAPTCHAs, you might need manual intervention or a CAPTCHA-solving service (though that adds complexity).

内容的提问来源于stack exchange,提问作者BJonas88

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 17:57:47