使用Selenium爬取dastelefonbuch.de5万+牙医数据 现有代码仅获19条
爬取德国电话黄页牙医数据的批量采集方案
现有代码问题排查
只能拿到19条数据的核心原因有2个:
- 仅请求了第一页搜索结果,站点默认每页仅展示20条左右商户信息,没有实现翻页/加载更多逻辑
- 提取电话的XPath用了全局路径
//span[@class='nr'],不是相对当前vcard的路径,会导致所有电话字段都取到第一条商户的电话,属于逻辑错误
批量采集实现方案
前置注意事项
- 该站点有反爬机制,频繁请求会触发验证码、IP封禁,建议采集间隔控制在2-5秒,必要时搭配代理IP池使用
- 站点总牙医商户数量确实超过5万,需要按分页逻辑逐页采集,总页数约2500+(单页20条)
优化后的代码实现
from selenium.webdriver.common.by import By from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import pandas as pd import random # 初始化存储列表 data_list = [] # chromedriver路径 PATH = r"C:\Program Files (x86)\chromedriver.exe" # 基础搜索url base_url = 'https://www.dastelefonbuch.de/Suche/Zahnarzt' # 初始化浏览器 options = webdriver.ChromeOptions() # 可选:添加无头模式后台运行 # options.add_argument('--headless=new') options.add_argument('--disable-blink-features=AutomationControlled') options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(PATH, options=options) driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', { 'source': 'Object.defineProperty(navigator, "webdriver", {get: () => undefined})' }) # 采集最大页数,可根据需求调整,5万条需要设置为2500+ max_page = 100 current_page = 1 try: driver.get(base_url) # 处理cookie提示弹窗 try: cookie_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.ID, 'onetrust-accept-btn-handler')) ) cookie_btn.click() time.sleep(2) except: pass while current_page <= max_page: print(f"正在采集第{current_page}页数据") # 等待当前页所有vcard加载完成 WebDriverWait(driver, 15).until( EC.presence_of_all_elements_located((By.XPATH, "//div[@class='vcard']")) ) # 滚动到底部加载所有条目 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(random.uniform(1, 2)) vid = driver.find_elements(By.XPATH, "//div[@class='vcard']") for item in vid: try: title = item.find_element(By.XPATH, ".//div[@class='name']").text.strip() except: title = '' try: # 修正为相对路径取电话 phone = item.find_element(By.XPATH, ".//span[@class='nr']").text.strip() except: phone = '' try: website = item.find_element(By.XPATH, ".//div[@class='url']/a").get_attribute('href').strip() except: website = '' try: state = item.find_element(By.XPATH, ".//div[@class='category']").text.strip() except: state = '' try: address = item.find_element(By.XPATH, ".//a[@class='addr']").text.strip() except: address = '' # 去重:如果当前商户名称和电话都为空则跳过 if title or phone: data_list.append({ "title": title, "phone": phone, "website": website, "state": state, "address": address }) # 翻页逻辑:触发下一页按钮 try: next_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//a[@class='next']")) ) next_btn.click() current_page += 1 # 随机等待避免被反爬识别 time.sleep(random.uniform(2, 4)) except: print("已到最后一页,采集结束") break finally: driver.quit() # 导出数据 df = pd.DataFrame(data_list) print(f"共采集到{len(df)}条数据") df.to_csv("牙医商户数据.csv", index=False, encoding='utf-8-sig')
额外优化建议
- 如果触发验证码,可以接入打码平台自动处理,或者手动设置等待时间人工验证
- 采集过程中可以每采集100页就做一次数据备份,避免程序崩溃导致数据丢失
- 若Selenium采集效率太低,可以抓站的请求接口直接用requests请求,接口返回JSON格式数据,采集效率会提升10倍以上
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

