使用Selenium爬取OpenSea时的两类问题求助:WebElement错误与无法获取超过6条结果
Let's break down and solve your two problems one by one:
Problem 1: selenium.webdriver.remote.webelement.webelement Error
Root Cause
The error pops up because you're trying to feed raw Selenium WebElement objects (like coll_name, art_name) directly into your Collection_Azuki dictionary, instead of using the text lists you spent time collecting (collection_name, name, etc.). Pandas can't serialize WebElement objects into a DataFrame, which triggers this messy error.
Fix
Update your dictionary to use the text lists you created, not the WebElement lists. I also fixed a small typo in your key name for consistency:
Collection_Azuki = { 'Collection_Name': collection_name, 'Collection_Description1': collection_desc1, 'Collection_Des2': collection_desc2, 'Collection_Des3': collection_desc3, # Fixed typo from "Colltionec_Des3" 'Art_Name_fav': name, 'Art_Price_fav': price }
Problem 2: Only Getting 6 Collection Results
Root Cause
OpenSea loads items dynamically as you scroll down. By default, it only loads 6 items initially, then waits for you to scroll to load more. Your current code doesn't handle this lazy loading, so it only captures the first batch of items.
Fix
Add a scroll loop to simulate scrolling to the bottom repeatedly until no new items load. We'll also replace unreliable time.sleep() calls with explicit waits to make the code more stable (page load times can vary, so hardcoded sleeps often fail).
Here's the full updated code with both fixes:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.service import Service import pandas as pd import time website = 'https://opensea.io/collection/azuki?search[sortAscending]=false&search[sortBy]=FAVORITE_COUNT' s = Service('C:\\webdrivers\\chromedriver.exe') # Fixed path with double backslashes for Windows driver = webdriver.Chrome(service=s) driver.get(website) driver.maximize_window() # Wait for "See more" button to be clickable, then expand description wait = WebDriverWait(driver, 10) see_more = wait.until(EC.element_to_be_clickable((By.XPATH, '//div[contains(text(), "See more")]'))) see_more.click() # Scroll to load all available items last_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll to bottom of page driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new content to load time.sleep(2) # Check if new content loaded new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break # No more items to load last_height = new_height # Collect elements after all content is loaded coll_name = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//h1'))) coll_desc1 = wait.until(EC.presence_of_all_elements_located((By.XPATH, '(//p[1])[1]'))) coll_desc2 = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//p[2]'))) coll_desc3 = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//p[3]'))) art_name = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//div[contains(@class, "sc-7qr9y8-0 sc-dw611d-1")]'))) art_price = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//div[contains(@class, "Price--amount")]'))) # Simplify text collection with list comprehensions collection_name = [c.text for c in coll_name] collection_desc1 = [d1.text for d1 in coll_desc1] collection_desc2 = [d2.text for d2 in coll_desc2] collection_desc3 = [d3.text for d3 in coll_desc3] name = [n.text for n in art_name] price = [p.text for p in art_price] # Build DataFrame with clean text data Collection_Azuki = { 'Collection_Name': collection_name, 'Collection_Description1': collection_desc1, 'Collection_Des2': collection_desc2, 'Collection_Des3': collection_desc3, 'Art_Name_fav': name, 'Art_Price_fav': price } df = pd.DataFrame.from_dict(Collection_Azuki, orient='index').transpose() print(df) driver.quit()
Key Improvements:
- Explicit Waits: Uses
WebDriverWaitto wait for elements to load properly, avoiding race conditions thattime.sleep()can miss. - Scroll Loop: Continuously loads new items until no more content is available.
- Robust XPath: Uses
contains(@class, ...)instead of exact class names, since OpenSea's dynamic class names might change over time. - Cleaner Code: Replaced repetitive for loops with list comprehensions for readability.
内容的提问来源于stack exchange,提问作者Muhammad Tayyab Shah

