Selenium爬取cci.fr房产中介数据时分页下一页点击失效求助
问题定位
- 致命运行错误:代码中
number_of_pages = int(driver.find_element_by_xpath('//span[contains(text(),"suivant")]'))行完全无法执行,一是find_element_by_xpath返回的是WebElement元素对象,不能直接转为int类型,运行到此处会直接抛出类型异常,后续下一页点击逻辑根本不会被触发;二是HTML中下一页按钮的文字是大写开头的Suivant,小写的匹配规则也找不到对应元素。 - 下一页点击稳定性不足:直接调用元素
click()方法容易遇到元素未加载完成、不在可视区域被遮挡的问题,导致点击无响应。 - 冗余代码干扰:你传入的初始URL已经是带搜索参数的结果页,不需要再执行搜索框输入、点击搜索按钮的冗余代码,多余的sleep也会增加运行不稳定概率。
解决方案
核心修改点
- 删除错误的页码统计行和冗余的搜索相关代码
- 引入显式等待确保下一页按钮加载完成后再操作
- 改用JS执行点击,避免元素被遮挡导致点击失效
- 统一CSV表头和内容的分隔符
修正后完整代码
from selenium import webdriver from selenium.common.exceptions import NoSuchElementException, TimeoutException from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import time # 初始化CSV文件,表头和写入分隔符统一用分号 with open('scraping_5_pagination.csv', 'w', encoding='utf-8') as file: file.write("business_names;attestation;town_pc;region\n") # 初始化浏览器 driver = webdriver.Chrome(ChromeDriverManager().install()) driver.get( 'https://www.cci.fr/agent-immobilier?company_name=agences%20immobili%C3%A8res%20&brand_name=&siren=&numero_carte=&code_region=84&city=&code_postal=&person_name=&state_recherche=1&name_region=AUVERGNE-RHONE-ALPES') driver.maximize_window() time.sleep(3) # 分页循环 for i in range(200): # 提取当前页数据 business_names = driver.find_elements(By.XPATH, '//td[@class="titre_entreprise"]') attestation = driver.find_elements(By.XPATH, '//tr[@class="lien-fiche"]/td/a') town_pc = driver.find_elements(By.XPATH, '//*[@id="main-content"]/div/table/tbody/tr/td[2]') region = driver.find_elements(By.XPATH, '//*[@id="main-content"]/div/table/tbody/tr/td[3]') # 写入CSV with open('scraping_5_pagination.csv', 'a', encoding='utf-8') as file: for j in range(len(business_names)): file.write( business_names[j].text + ";" + attestation[j].text + ";" + town_pc[j].text + ";" + region[j].text + "\n") # 尝试点击下一页 try: # 显式等待下一页按钮可点击 next_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//a[@rel="next"]')) ) # 用JS点击避免元素遮挡问题 driver.execute_script("arguments[0].click();", next_btn) time.sleep(2) except (NoSuchElementException, TimeoutException): # 没有下一页时退出循环 break driver.quit()
内容的提问来源于stack exchange,提问作者radia12
相关产品推荐
相关产品推荐

