You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium实现Doctolib自动加载更多并抓取表格存CSV?

解决方案:Doctolib巴黎全科医生数据抓取(自动加载全部结果)

原代码存在的问题

  • 重复导入相同模块,代码冗余
  • 使用绝对XPath定位按钮,页面结构稍有变化就会失效
  • 仅执行一次按钮点击,未实现循环加载直到无更多结果
  • 全局变量driver被重复初始化,逻辑冲突
  • 固定time.sleep等待不够灵活,易出现元素未加载完成就查找的错误
  • 未处理元素不存在的异常场景

修正后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, TimeoutException
import pandas as pd
from bs4 import BeautifulSoup

def init_driver():
    # 初始化Chrome浏览器配置
    options = Options()
    options.add_argument("--disable-infobars")
    options.add_argument("--start-maximized")
    # 自动管理驱动,无需手动指定本地路径
    driver = webdriver.Chrome(options=options)
    return driver

def load_all_results(driver, url):
    driver.get(url)
    wait = WebDriverWait(driver, 10)
    
    while True:
        try:
            # 用相对XPath定位加载按钮,适配页面结构变化
            load_more_btn = wait.until(
                EC.element_to_be_clickable(
                    (By.XPATH, "//button[contains(text(), 'afficher plus de résultats')]")
                )
            )
            load_more_btn.click()
            # 等待新内容加载完成,避免重复点击
            wait.until(EC.staleness_of(load_more_btn))
        except (NoSuchElementException, TimeoutException):
            # 按钮不存在或超时,说明已加载全部结果
            break

def extract_doctor_data(driver):
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    doctors = []
    
    # 遍历所有医生卡片
    for card in soup.find_all('div', class_='dl-search-result'):
        # 提取姓名
        name_tag = card.find('h3', class_='dl-search-result-name')
        name = name_tag.get_text(strip=True) if name_tag else None
        
        # 提取地址
        address_tag = card.find('div', class_='dl-search-result-address')
        address = address_tag.get_text(strip=True) if address_tag else None
        
        # 提取专科信息(可按需扩展其他字段)
        specialty_tag = card.find('div', class_='dl-search-result-specialty')
        specialty = specialty_tag.get_text(strip=True) if specialty_tag else None
        
        doctors.append({
            '姓名': name,
            '地址': address,
            '专科': specialty
        })
    
    return pd.DataFrame(doctors)

if __name__ == "__main__":
    target_url = "https://www.doctolib.fr/medecin-generaliste/paris?availabilities=3"
    driver = init_driver()
    
    try:
        load_all_results(driver, target_url)
        doctor_df = extract_doctor_data(driver)
        # 保存数据到CSV文件
        doctor_df.to_csv('巴黎全科医生数据.csv', index=False, encoding='utf-8-sig')
        print(f"数据抓取完成,共获取{len(doctor_df)}条记录")
    finally:
        driver.quit()

代码说明

  • 驱动管理:自动适配Chrome驱动,无需手动配置本地路径,兼容不同系统
  • 循环加载:通过显式等待定位加载按钮,点击后等待页面更新,直到按钮消失自动终止循环
  • 定位优化:使用包含按钮文本的相对XPath,避免绝对路径的不稳定性
  • 数据提取:用BeautifulSoup解析最终页面,提取核心字段并保存为CSV文件
  • 异常处理:捕获元素不存在和超时异常,确保程序稳定运行

内容的提问来源于stack exchange,提问作者TMB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:55:23