如何用Selenium循环爬取OpenML平台的多页数据?
Solution to Scrape All OpenML Task Runs Data
Got it, let's fix this so you can scrape all 400k+ results from OpenML's task runs page. The main issue with your current code is that you're only clicking the next page once—we need a loop that keeps going until there are no more pages left. Here's how to do it properly:
from selenium import webdriver import time # Initialize the Chrome driver chrome_path = r"C:\Users\Zeshan\Desktop\chromedriver_win32\chromedriver.exe" driver = webdriver.Chrome(chrome_path) driver.get("https://www.openml.org/t/31#!taskruns") # Track all scraped items across pages total_items = 0 try: while True: # Wait for page to fully load (adjust sleep time based on your internet speed) time.sleep(3) # Scrape current page's data titles = driver.find_elements_by_xpath('//div[@class="itemheadfull"]') metrics = driver.find_elements_by_xpath('//div[@class="runStats statLine"]') # Process and print each item for i in range(len(titles)): total_items += 1 print(f"{titles[i].text} + {metrics[i].text}") print(f"Output Number: {total_items}") # Attempt to navigate to the next page try: # Target the "next results now.." link specifically next_page = driver.find_element_by_xpath('//*[@id="taskruns"]/div/p/a[contains(text(), "next results now..")]') # Check if the link is clickable (not disabled, meaning there are more pages) if next_page.is_enabled(): next_page.click() # Wait for the loading spinner to finish time.sleep(2) # Extra check: Wait until "Loading more..." text disappears while driver.find_elements_by_xpath('//div[contains(text(), "Loading more...")]'): time.sleep(1) else: # No more pages available, exit loop break except: # If next page link isn't found, we've reached the end of results print("Finished scraping all pages!") break finally: # Clean up: Close the browser when done driver.quit()
Key Improvements:
- Continuous Loop: The
while Trueloop runs indefinitely until we hit the last page, ensuring we scrape every result set. - Load Wait Logic: Added sleep timers and a check for the "Loading more..." state to avoid scraping incomplete page data.
- Total Item Tracking: A single counter keeps track of all scraped items across pages, so you don't have to manually offset numbering per page.
- Error Handling: The inner
try/exceptblock catches cases where the next page link doesn't exist, gracefully ending the scrape when we're done.
Quick Notes:
- Adjust the
time.sleep()values if pages load slower for you—longer waits reduce the chance of missing data. - For 400k+ items, consider saving results to a CSV or database instead of printing, to avoid losing data if the script is interrupted.
- Be respectful of OpenML's servers—you might want to add longer delays between page clicks if you're scraping at scale.
内容的提问来源于stack exchange,提问作者Zeshan Fayyaz
相关产品推荐
相关产品推荐

