如何解析Etherscan?如何抓取指定代币持仓页所有ETH地址并存为TXT
Hey there! Let's tackle your question with two clear sections: first, how Etherscan parsing works, then a step-by-step solution to scrape those token holder addresses and save them to a text file.
Etherscan is a blockchain explorer that aggregates data from Ethereum nodes and stores it in its own database. When it comes to "parsing" Etherscan, you have two main approaches:
Official Etherscan API (Recommended)
This is the most reliable and compliant way to get data. Etherscan offers a free-tier API (with rate limits) that lets you fetch account balances, token holders, transaction histories, and more. You'll need to sign up for an API key on their site, but it's straightforward. Using the API avoids anti-scraping blocks and ensures you get structured, clean data (usually JSON format).Web Scraping (Use with Caution)
If you need data that's not available via the API, you can scrape the frontend HTML. But be aware: Etherscan has anti-scraping measures (like IP bans for frequent requests) and their terms of service restrict automated scraping. Always check theirrobots.txtfile and limit request frequency to avoid getting blocked.
For your specific target page (the token holder list), here's a Python-based solution that handles pagination and saves addresses to a text file.
Prerequisites
First, install the required libraries:
pip install requests beautifulsoup4 fake-useragent
Full Code Example
import requests from bs4 import BeautifulSoup from fake_useragent import UserAgent import time def scrape_etherscan_token_holders(contract_address, output_file="token_holders.txt"): ua = UserAgent() base_url = f"https://etherscan.io/token/generic-tokenholders2?a={contract_address}&s=0&p=" holders = set() # Use a set to avoid duplicate addresses page_num = 1 while True: # Generate random user agent to mimic a browser headers = {"User-Agent": ua.random} url = base_url + str(page_num) try: response = requests.get(url, headers=headers) response.raise_for_status() # Raise error for HTTP issues except requests.exceptions.RequestException as e: print(f"Request failed for page {page_num}: {e}") break soup = BeautifulSoup(response.text, "html.parser") # Find all address elements (adjust the selector if Etherscan updates their HTML) address_elements = soup.select("td.text-secondary a") if not address_elements: print("No more addresses found. Stopping scrape.") break # Extract addresses and add to the set for elem in address_elements: address = elem.get_text(strip=True) if address.startswith("0x"): holders.add(address) print(f"Scraped page {page_num}, total unique holders so far: {len(holders)}") page_num += 1 time.sleep(1.5) # Add delay to avoid triggering anti-scraping # Save to text file with open(output_file, "w") as f: for addr in holders: f.write(f"{addr}\n") print(f"Successfully saved {len(holders)} unique addresses to {output_file}") # Run the function with your contract address scrape_etherscan_token_holders("0x6425c6be902d692ae2db752b3c268afadb099d3b")
Key Notes
- HTML Selector: The selector
td.text-secondary atargets the address links in the holder table. If Etherscan updates their page structure, you'll need to inspect the HTML and adjust this selector. - Anti-Scraping Mitigation: Using
fake-useragentgenerates random browser headers, and thetime.sleep(1.5)adds a delay between requests to avoid being flagged as a bot. - Duplicate Prevention: Using a
setensures we don't save duplicate addresses. - API Alternative: For a more stable solution, use Etherscan's Token Holder API endpoint (requires an API key):
This returns JSON data, so you can skip HTML parsing entirely.https://api.etherscan.io/api?module=token&action=tokenholderlist&contractaddress=0x6425c6be902d692ae2db752b3c268afadb099d3b&page=1&offset=100&apikey=YOUR_API_KEY
内容的提问来源于stack exchange,提问作者Tref P

