You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium循环爬取OpenML平台的多页数据?

Solution to Scrape All OpenML Task Runs Data

Got it, let's fix this so you can scrape all 400k+ results from OpenML's task runs page. The main issue with your current code is that you're only clicking the next page once—we need a loop that keeps going until there are no more pages left. Here's how to do it properly:

from selenium import webdriver
import time

# Initialize the Chrome driver
chrome_path = r"C:\Users\Zeshan\Desktop\chromedriver_win32\chromedriver.exe"
driver = webdriver.Chrome(chrome_path)
driver.get("https://www.openml.org/t/31#!taskruns")

# Track all scraped items across pages
total_items = 0

try:
    while True:
        # Wait for page to fully load (adjust sleep time based on your internet speed)
        time.sleep(3)
        
        # Scrape current page's data
        titles = driver.find_elements_by_xpath('//div[@class="itemheadfull"]')
        metrics = driver.find_elements_by_xpath('//div[@class="runStats statLine"]')
        
        # Process and print each item
        for i in range(len(titles)):
            total_items += 1
            print(f"{titles[i].text} + {metrics[i].text}")
            print(f"Output Number: {total_items}")
        
        # Attempt to navigate to the next page
        try:
            # Target the "next results now.." link specifically
            next_page = driver.find_element_by_xpath('//*[@id="taskruns"]/div/p/a[contains(text(), "next results now..")]')
            
            # Check if the link is clickable (not disabled, meaning there are more pages)
            if next_page.is_enabled():
                next_page.click()
                # Wait for the loading spinner to finish
                time.sleep(2)
                # Extra check: Wait until "Loading more..." text disappears
                while driver.find_elements_by_xpath('//div[contains(text(), "Loading more...")]'):
                    time.sleep(1)
            else:
                # No more pages available, exit loop
                break
        except:
            # If next page link isn't found, we've reached the end of results
            print("Finished scraping all pages!")
            break

finally:
    # Clean up: Close the browser when done
    driver.quit()

Key Improvements:

  • Continuous Loop: The while True loop runs indefinitely until we hit the last page, ensuring we scrape every result set.
  • Load Wait Logic: Added sleep timers and a check for the "Loading more..." state to avoid scraping incomplete page data.
  • Total Item Tracking: A single counter keeps track of all scraped items across pages, so you don't have to manually offset numbering per page.
  • Error Handling: The inner try/except block catches cases where the next page link doesn't exist, gracefully ending the scrape when we're done.

Quick Notes:

  • Adjust the time.sleep() values if pages load slower for you—longer waits reduce the chance of missing data.
  • For 400k+ items, consider saving results to a CSV or database instead of printing, to avoid losing data if the script is interrupted.
  • Be respectful of OpenML's servers—you might want to add longer delays between page clicks if you're scraping at scale.

内容的提问来源于stack exchange,提问作者Zeshan Fayyaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:53:18