Python Selenium爬虫循环输出重复内容问题排查请求
问题排查:Selenium爬虫循环重复抓取同一诊所数据
核心问题原因
- XPath路径错误:在子元素查找时使用
//开头的绝对路径,会从整个HTML文档根节点开始匹配元素,而非当前诊所区块(blocks[block])的子节点,导致每次都抓取页面第一个诊所的数据。 - 循环逻辑不合理:固定循环30次,未考虑单页实际诊所数量,分页后未调整循环范围,易引发索引越界或重复数据问题。
修复后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd # 用原始字符串避免路径转义问题 s = Service(r"C:\selenium driver\chromedriver.exe") driver = webdriver.Chrome(service=s) companies_names = [] persons_names = [] phones_numbers = [] locations = [] opening_hours = [] descriptions = [] websites_links = [] all_profiles = [] driver.get("https://www.saveface.co.uk/search/") # 显式等待比隐式等待更可靠 wait = WebDriverWait(driver, 10) while True: # 等待当前页诊所区块加载完成 blocks = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//div[@class='result clientresult']"))) # 遍历当前页所有诊所区块 for block in blocks: # 使用相对XPath(.//)限定在当前block节点内查找子元素 company_name = block.find_element(By.XPATH, ".//h3[@class='resulttitle']").text.strip() companies_names.append(company_name) person_name = block.find_element(By.XPATH, ".//p[@class='name_wrapper']").text.strip() persons_names.append(person_name) phone_number = block.find_element(By.XPATH, ".//div[@class='searchContact phone']").text.strip() phones_numbers.append(phone_number) location = block.find_element(By.XPATH, ".//li[@class='cls_loc']").text.strip() locations.append(location) opening_hour = block.find_element(By.XPATH, ".//li[@class='opening-hours']").text.strip() opening_hours.append(opening_hour) profile = block.find_element(By.XPATH, ".//a[@class='visitpage']").get_attribute("href") all_profiles.append(profile) print(company_name, person_name, phone_number, location, opening_hour, profile) # 尝试查找下一页按钮,无则退出循环 try: next_page = wait.until(EC.element_to_be_clickable((By.XPATH, "//a[@class='facetwp-page' and not(@class='facetwp-page active')]"))) next_page.click() # 等待页面刷新完成 wait.until(EC.staleness_of(blocks[0])) except: print("已无更多页面,停止抓取") break # 抓取诊所详情页数据 for profile_url in all_profiles: driver.get(profile_url) try: description = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='desc-text-left']"))).text.strip() descriptions.append(description) except: descriptions.append("无描述信息") try: website_link = wait.until(EC.presence_of_element_located((By.XPATH, "//a[@class='visitwebsite website']"))).get_attribute("href") websites_links.append(website_link) except: websites_links.append("无官网链接") driver.close() # 生成CSV文件 df = pd.DataFrame( { "company_name": companies_names, "person_name": persons_names, "phone_number": phones_numbers, "location": locations, "opening_hour": opening_hours, "description": descriptions, "website_link": websites_links, "profile_on_saveface": all_profiles } ) df.to_csv('saveface.csv', index=False)
关键修复点说明
- 相对XPath路径:将子元素查找的XPath改为
.//开头,限定在当前诊所区块内查找,确保获取对应诊所的数据。 - 动态分页循环:用
while True自动处理分页,直到找不到下一页按钮,适配不同页面的诊所数量。 - 显式等待优化:使用
WebDriverWait等待元素加载,避免页面加载慢导致的元素未找到问题。 - 异常捕获:详情页抓取时添加异常处理,避免因部分诊所缺少信息导致程序崩溃。
- 路径转义处理:用原始字符串定义chromedriver路径,避免反斜杠转义错误。
内容的提问来源于stack exchange,提问作者Zeyad Magdy
相关产品推荐
相关产品推荐

