You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬虫循环输出重复内容问题排查请求

问题排查:Selenium爬虫循环重复抓取同一诊所数据

核心问题原因

  1. XPath路径错误:在子元素查找时使用//开头的绝对路径,会从整个HTML文档根节点开始匹配元素,而非当前诊所区块(blocks[block])的子节点,导致每次都抓取页面第一个诊所的数据。
  2. 循环逻辑不合理:固定循环30次,未考虑单页实际诊所数量,分页后未调整循环范围,易引发索引越界或重复数据问题。

修复后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

# 用原始字符串避免路径转义问题
s = Service(r"C:\selenium driver\chromedriver.exe")
driver = webdriver.Chrome(service=s)

companies_names = []
persons_names = []
phones_numbers = []
locations = []
opening_hours = []
descriptions = []
websites_links = []
all_profiles = []

driver.get("https://www.saveface.co.uk/search/")
# 显式等待比隐式等待更可靠
wait = WebDriverWait(driver, 10)

while True:
    # 等待当前页诊所区块加载完成
    blocks = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//div[@class='result clientresult']")))
    
    # 遍历当前页所有诊所区块
    for block in blocks:
        # 使用相对XPath(.//)限定在当前block节点内查找子元素
        company_name = block.find_element(By.XPATH, ".//h3[@class='resulttitle']").text.strip()
        companies_names.append(company_name)

        person_name = block.find_element(By.XPATH, ".//p[@class='name_wrapper']").text.strip()
        persons_names.append(person_name)

        phone_number = block.find_element(By.XPATH, ".//div[@class='searchContact phone']").text.strip()
        phones_numbers.append(phone_number)

        location = block.find_element(By.XPATH, ".//li[@class='cls_loc']").text.strip()
        locations.append(location)

        opening_hour = block.find_element(By.XPATH, ".//li[@class='opening-hours']").text.strip()
        opening_hours.append(opening_hour)

        profile = block.find_element(By.XPATH, ".//a[@class='visitpage']").get_attribute("href")
        all_profiles.append(profile)
        
        print(company_name, person_name, phone_number, location, opening_hour, profile)
    
    # 尝试查找下一页按钮,无则退出循环
    try:
        next_page = wait.until(EC.element_to_be_clickable((By.XPATH, "//a[@class='facetwp-page' and not(@class='facetwp-page active')]")))
        next_page.click()
        # 等待页面刷新完成
        wait.until(EC.staleness_of(blocks[0]))
    except:
        print("已无更多页面,停止抓取")
        break

# 抓取诊所详情页数据
for profile_url in all_profiles:
    driver.get(profile_url)
    try:
        description = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='desc-text-left']"))).text.strip()
        descriptions.append(description)
    except:
        descriptions.append("无描述信息")
    
    try:
        website_link = wait.until(EC.presence_of_element_located((By.XPATH, "//a[@class='visitwebsite website']"))).get_attribute("href")
        websites_links.append(website_link)
    except:
        websites_links.append("无官网链接")

driver.close()

# 生成CSV文件
df = pd.DataFrame(
    {
        "company_name": companies_names,
        "person_name": persons_names,
        "phone_number": phones_numbers,
        "location": locations,
        "opening_hour": opening_hours,
        "description": descriptions,
        "website_link": websites_links,
        "profile_on_saveface": all_profiles
    }
)

df.to_csv('saveface.csv', index=False)

关键修复点说明

  • 相对XPath路径:将子元素查找的XPath改为.//开头,限定在当前诊所区块内查找,确保获取对应诊所的数据。
  • 动态分页循环:用while True自动处理分页,直到找不到下一页按钮,适配不同页面的诊所数量。
  • 显式等待优化:使用WebDriverWait等待元素加载,避免页面加载慢导致的元素未找到问题。
  • 异常捕获:详情页抓取时添加异常处理,避免因部分诊所缺少信息导致程序崩溃。
  • 路径转义处理:用原始字符串定义chromedriver路径,避免反斜杠转义错误。

内容的提问来源于stack exchange,提问作者Zeyad Magdy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 18:05:23