You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取dastelefonbuch.de5万+牙医数据 现有代码仅获19条

爬取德国电话黄页牙医数据的批量采集方案

现有代码问题排查

只能拿到19条数据的核心原因有2个:

  1. 仅请求了第一页搜索结果,站点默认每页仅展示20条左右商户信息,没有实现翻页/加载更多逻辑
  2. 提取电话的XPath用了全局路径//span[@class='nr'],不是相对当前vcard的路径,会导致所有电话字段都取到第一条商户的电话,属于逻辑错误

批量采集实现方案

前置注意事项

  • 该站点有反爬机制,频繁请求会触发验证码、IP封禁,建议采集间隔控制在2-5秒,必要时搭配代理IP池使用
  • 站点总牙医商户数量确实超过5万,需要按分页逻辑逐页采集,总页数约2500+(单页20条)

优化后的代码实现

from selenium.webdriver.common.by import By
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import pandas as pd
import random

# 初始化存储列表
data_list = []
# chromedriver路径
PATH = r"C:\Program Files (x86)\chromedriver.exe"
# 基础搜索url
base_url = 'https://www.dastelefonbuch.de/Suche/Zahnarzt'
# 初始化浏览器
options = webdriver.ChromeOptions()
# 可选:添加无头模式后台运行
# options.add_argument('--headless=new')
options.add_argument('--disable-blink-features=AutomationControlled')
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
driver = webdriver.Chrome(PATH, options=options)
driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', {
    'source': 'Object.defineProperty(navigator, "webdriver", {get: () => undefined})'
})

# 采集最大页数,可根据需求调整,5万条需要设置为2500+
max_page = 100
current_page = 1

try:
    driver.get(base_url)
    # 处理cookie提示弹窗
    try:
        cookie_btn = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.ID, 'onetrust-accept-btn-handler'))
        )
        cookie_btn.click()
        time.sleep(2)
    except:
        pass

    while current_page <= max_page:
        print(f"正在采集第{current_page}页数据")
        # 等待当前页所有vcard加载完成
        WebDriverWait(driver, 15).until(
            EC.presence_of_all_elements_located((By.XPATH, "//div[@class='vcard']"))
        )
        # 滚动到底部加载所有条目
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(random.uniform(1, 2))
        
        vid = driver.find_elements(By.XPATH, "//div[@class='vcard']")
        for item in vid:
            try:
                title = item.find_element(By.XPATH, ".//div[@class='name']").text.strip()
            except:
                title = ''
            try:
                # 修正为相对路径取电话
                phone = item.find_element(By.XPATH, ".//span[@class='nr']").text.strip()
            except:
                phone = ''
            try:
                website = item.find_element(By.XPATH, ".//div[@class='url']/a").get_attribute('href').strip()
            except:
                website = ''
            try:
                state = item.find_element(By.XPATH, ".//div[@class='category']").text.strip()
            except:
                state = ''
            try:
                address = item.find_element(By.XPATH, ".//a[@class='addr']").text.strip()
            except:
                address = ''
            # 去重:如果当前商户名称和电话都为空则跳过
            if title or phone:
                data_list.append({
                    "title": title,
                    "phone": phone,
                    "website": website,
                    "state": state,
                    "address": address
                })
        
        # 翻页逻辑:触发下一页按钮
        try:
            next_btn = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.XPATH, "//a[@class='next']"))
            )
            next_btn.click()
            current_page += 1
            # 随机等待避免被反爬识别
            time.sleep(random.uniform(2, 4))
        except:
            print("已到最后一页,采集结束")
            break
finally:
    driver.quit()

# 导出数据
df = pd.DataFrame(data_list)
print(f"共采集到{len(df)}条数据")
df.to_csv("牙医商户数据.csv", index=False, encoding='utf-8-sig')

额外优化建议

  • 如果触发验证码,可以接入打码平台自动处理,或者手动设置等待时间人工验证
  • 采集过程中可以每采集100页就做一次数据备份,避免程序崩溃导致数据丢失
  • 若Selenium采集效率太低,可以抓站的请求接口直接用requests请求,接口返回JSON格式数据,采集效率会提升10倍以上

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 22:45:10