如何用Selenium Python循环点击下一页爬取所有页面企业名称?
Fixing Selenium Pagination for Scraping All Company Pages
Hey Neha, let's get that pagination loop working so you can scrape every company name across all pages! Your current code only grabs data from the first page, so we'll adjust it to keep clicking "Next" until there are no more pages left.
Modified Full Code
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException import pandas as pd DRIVER_PATH = 'C:/Users/Neha/Downloads/chromedriver_win32/chromedriver' driver = webdriver.Chrome(executable_path=DRIVER_PATH) driver.get('https://www.fundoodata.com/citiesindustry/19/2/list-of-information-technology-(it)-companies-in-noida') names_list = [] while True: # Scrape company names from current page company_names = driver.find_elements(By.CLASS_NAME, 'heading') for name in company_names: text = name.text.strip() if text: # Skip empty entries names_list.append(text) print(text) # Try to click "Next" button try: # Wait for the next button to be clickable (max 10 seconds) next_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="main-container"]/div[2]/div[4]/div[2]/div[1]/div/ul/li[7]/a')) ) next_button.click() # Wait for the next page to load (wait until company names appear) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'heading')) ) except (NoSuchElementException, ElementClickInterceptedException): # No more pages or button is unclickable, exit loop print("No more pages to scrape. Exiting loop.") break # Clean up and save data driver.quit() df = pd.DataFrame(names_list, columns=['Company Name']) # Add column name for clarity writer = pd.ExcelWriter('companies_names.xlsx', engine='xlsxwriter') df.to_excel(writer, sheet_name='List', index=False) # Hide index column in Excel writer.close()
Key Improvements Explained
- Pagination Loop: The
while Trueloop runs until we can't find or click the "Next" button anymore. - Reliable Waits: We use
WebDriverWaitinstead of hardcodedtime.sleep()to wait for elements to be ready—this makes the code more stable for varying page load speeds. - Exception Handling: We catch two common exceptions that signal the end of pagination:
NoSuchElementException: The "Next" button doesn't exist (we're on the last page).ElementClickInterceptedException: The button exists but is blocked (e.g., by a popup or overlay).
- Cleaner Data: We add
.strip()to remove extra whitespace and check for non-empty text to avoid adding blank entries to our list. - Modern Selenium Syntax: Replaced the deprecated
find_elements_by_class_namewith the recommendedfind_elements(By.CLASS_NAME, ...)syntax for Selenium 4+. - Better Excel Output: Added a column name (
Company Name) and hid the default index column to make the Excel file more readable.
内容的提问来源于stack exchange,提问作者Neha Sharma
相关产品推荐
相关产品推荐

