You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:提取https://www.tokendata.io/网站表格数据

How to Scrape Data from tokendata.io Using Selenium (For Beginners)

Hey there! I totally get how urgent this is for your thesis—dynamic sites like tokendata.io rely on JavaScript to load content, which is why BeautifulSoup alone can’t grab the data (it only parses static HTML). Let’s walk through exactly how to use Selenium to get what you need, step by step.

Step 1: Install Required Tools

First, make sure you have these set up:

  • Selenium: Run pip install selenium in your terminal
  • A browser driver (we’ll use ChromeDriver here, since it’s the most common):
    • Download the version that matches your Chrome browser (search for “ChromeDriver download” to find the official page)
    • Place the driver executable in a system PATH location, or note its file path for later

Step 2: Basic Selenium Setup to Load the Page

Here’s a simple script to launch Chrome, navigate to the site, and wait for the page to fully load (critical for dynamic content):

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
import pandas as pd

# Initialize the Chrome driver (replace the path if you didn't add it to PATH)
driver = webdriver.Chrome(executable_path="/path/to/chromedriver")  # Or just webdriver.Chrome() if in PATH

# Navigate to the target site
driver.get("https://www.tokendata.io/")

# Wait for the table to load (adjust timeout if needed—10 seconds is usually enough)
try:
    # Wait until the table element is present on the page
    table = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.TAG_NAME, "table"))
    )
except TimeoutException:
    print("Timed out waiting for the table to load")
    driver.quit()

Step 3: Extract Table Data

Once the table loads, we can pull all rows and cells:

# Grab all rows from the table
rows = table.find_elements(By.TAG_NAME, "tr")

# Initialize a list to store our scraped data
scraped_data = []

# Loop through each row to extract cell text
for row in rows:
    cells = row.find_elements(By.TAG_NAME, "td")
    # Clean up text and collect data for the row
    row_content = [cell.text.strip() for cell in cells]
    # Skip empty rows
    if row_content:
        scraped_data.append(row_content)

# Extract header text if the table has headers
header_cells = table.find_elements(By.TAG_NAME, "th")
table_headers = [header.text.strip() for header in header_cells]

Step 4: Convert to a Usable Format (Pandas DataFrame)

For your thesis, you’ll want structured data you can analyze or export—Pandas makes this easy:

# Create a DataFrame with headers and scraped data
df = pd.DataFrame(scraped_data, columns=table_headers)

# Save to a CSV file (perfect for thesis analysis!)
df.to_csv("tokendata_scraped.csv", index=False)
print("Data saved to tokendata_scraped.csv successfully!")

Step 5: Clean Up

Don’t forget to close the browser when you’re done:

driver.quit()

Key Tips for Your Thesis

  • Be respectful: Don’t run the script too frequently—this can get your IP blocked. Add a small delay (like time.sleep(2) after importing time) if scraping multiple pages.
  • Headless mode: If you don’t want the Chrome window to pop up every time, run the driver in headless mode:
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")
    driver = webdriver.Chrome(options=options)
    
  • Scroll handling: If the table loads more content as you scroll, let me know—I can add code to handle that!

I hope this gets you the data you need quickly—good luck with your thesis!

内容的提问来源于stack exchange,提问作者Alberto Giro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:14:47