You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium Python循环点击下一页爬取所有页面企业名称?

Fixing Selenium Pagination for Scraping All Company Pages

Hey Neha, let's get that pagination loop working so you can scrape every company name across all pages! Your current code only grabs data from the first page, so we'll adjust it to keep clicking "Next" until there are no more pages left.

Modified Full Code

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException
import pandas as pd

DRIVER_PATH = 'C:/Users/Neha/Downloads/chromedriver_win32/chromedriver'
driver = webdriver.Chrome(executable_path=DRIVER_PATH)
driver.get('https://www.fundoodata.com/citiesindustry/19/2/list-of-information-technology-(it)-companies-in-noida')

names_list = []

while True:
    # Scrape company names from current page
    company_names = driver.find_elements(By.CLASS_NAME, 'heading')
    for name in company_names:
        text = name.text.strip()
        if text:  # Skip empty entries
            names_list.append(text)
            print(text)
    
    # Try to click "Next" button
    try:
        # Wait for the next button to be clickable (max 10 seconds)
        next_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, '//*[@id="main-container"]/div[2]/div[4]/div[2]/div[1]/div/ul/li[7]/a'))
        )
        next_button.click()
        
        # Wait for the next page to load (wait until company names appear)
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CLASS_NAME, 'heading'))
        )
    except (NoSuchElementException, ElementClickInterceptedException):
        # No more pages or button is unclickable, exit loop
        print("No more pages to scrape. Exiting loop.")
        break

# Clean up and save data
driver.quit()

df = pd.DataFrame(names_list, columns=['Company Name'])  # Add column name for clarity
writer = pd.ExcelWriter('companies_names.xlsx', engine='xlsxwriter')
df.to_excel(writer, sheet_name='List', index=False)  # Hide index column in Excel
writer.close()

Key Improvements Explained

  • Pagination Loop: The while True loop runs until we can't find or click the "Next" button anymore.
  • Reliable Waits: We use WebDriverWait instead of hardcoded time.sleep() to wait for elements to be ready—this makes the code more stable for varying page load speeds.
  • Exception Handling: We catch two common exceptions that signal the end of pagination:
    • NoSuchElementException: The "Next" button doesn't exist (we're on the last page).
    • ElementClickInterceptedException: The button exists but is blocked (e.g., by a popup or overlay).
  • Cleaner Data: We add .strip() to remove extra whitespace and check for non-empty text to avoid adding blank entries to our list.
  • Modern Selenium Syntax: Replaced the deprecated find_elements_by_class_name with the recommended find_elements(By.CLASS_NAME, ...) syntax for Selenium 4+.
  • Better Excel Output: Added a column name (Company Name) and hid the default index column to make the Excel file more readable.

内容的提问来源于stack exchange,提问作者Neha Sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 20:17:37