Python新手求助:提取https://www.tokendata.io/网站表格数据
Hey there! I totally get how urgent this is for your thesis—dynamic sites like tokendata.io rely on JavaScript to load content, which is why BeautifulSoup alone can’t grab the data (it only parses static HTML). Let’s walk through exactly how to use Selenium to get what you need, step by step.
Step 1: Install Required Tools
First, make sure you have these set up:
- Selenium: Run
pip install seleniumin your terminal - A browser driver (we’ll use ChromeDriver here, since it’s the most common):
- Download the version that matches your Chrome browser (search for “ChromeDriver download” to find the official page)
- Place the driver executable in a system PATH location, or note its file path for later
Step 2: Basic Selenium Setup to Load the Page
Here’s a simple script to launch Chrome, navigate to the site, and wait for the page to fully load (critical for dynamic content):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException import pandas as pd # Initialize the Chrome driver (replace the path if you didn't add it to PATH) driver = webdriver.Chrome(executable_path="/path/to/chromedriver") # Or just webdriver.Chrome() if in PATH # Navigate to the target site driver.get("https://www.tokendata.io/") # Wait for the table to load (adjust timeout if needed—10 seconds is usually enough) try: # Wait until the table element is present on the page table = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "table")) ) except TimeoutException: print("Timed out waiting for the table to load") driver.quit()
Step 3: Extract Table Data
Once the table loads, we can pull all rows and cells:
# Grab all rows from the table rows = table.find_elements(By.TAG_NAME, "tr") # Initialize a list to store our scraped data scraped_data = [] # Loop through each row to extract cell text for row in rows: cells = row.find_elements(By.TAG_NAME, "td") # Clean up text and collect data for the row row_content = [cell.text.strip() for cell in cells] # Skip empty rows if row_content: scraped_data.append(row_content) # Extract header text if the table has headers header_cells = table.find_elements(By.TAG_NAME, "th") table_headers = [header.text.strip() for header in header_cells]
Step 4: Convert to a Usable Format (Pandas DataFrame)
For your thesis, you’ll want structured data you can analyze or export—Pandas makes this easy:
# Create a DataFrame with headers and scraped data df = pd.DataFrame(scraped_data, columns=table_headers) # Save to a CSV file (perfect for thesis analysis!) df.to_csv("tokendata_scraped.csv", index=False) print("Data saved to tokendata_scraped.csv successfully!")
Step 5: Clean Up
Don’t forget to close the browser when you’re done:
driver.quit()
Key Tips for Your Thesis
- Be respectful: Don’t run the script too frequently—this can get your IP blocked. Add a small delay (like
time.sleep(2)after importingtime) if scraping multiple pages. - Headless mode: If you don’t want the Chrome window to pop up every time, run the driver in headless mode:
options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(options=options) - Scroll handling: If the table loads more content as you scroll, let me know—I can add code to handle that!
I hope this gets you the data you need quickly—good luck with your thesis!
内容的提问来源于stack exchange,提问作者Alberto Giro

