You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无动态URL的动态网站爬取:牙科诊所数据爬取方案求助

牙科诊所数据爬取问题解决方案

我需要爬取所有牙科诊所的数据,目前有两个代码实现方案但均存在问题:

  1. 方案1已实现单页诊所信息爬取,但不知如何设置循环实现多页爬取
  2. 方案2通过API请求获取数据,但不知如何提取地址、电话、邮箱等其余数据

方案1:单页爬取扩展为多页爬取

目标网站分页通过page参数实现(如第2页URL为https://www.dent.cz/zubni-lekari?page=2),以下是修改后的多页爬取代码:

from requests_html import HTMLSession

base_url = "https://www.dent.cz/zubni-lekari"
s = HTMLSession()

# 获取总页数,找不到则兜底设为5页,可按需调整
r = s.get(base_url)
r.html.render(sleep=1)
last_page_elem = r.html.xpath('//*[@class="pagination"]//a[last()-1]/text()', first=True)
total_pages = int(last_page_elem.text) if last_page_elem else 5

all_clinics = []

for page in range(1, total_pages + 1):
    url = f"{base_url}?page={page}"
    r = s.get(url)
    r.html.render(sleep=1)
    
    # 遍历当前页的诊所,最多30个
    for x in range(1, 31):
        clinic = r.html.xpath(f'//*[@id="main"]/div/div[3]/div[1]/div/div[{x}]/h3', first=True)
        if not clinic:  # 无数据则跳出当前页循环
            break
        
        adress = r.html.xpath(f'//*[@id="main"]/div/div[3]/div[1]/div/div[{x}]/p[1]', first=True)
        phone = r.html.xpath(f'//*[@id="main"]/div/div[3]/div[1]/div/div[{x}]/p[1]/strong[1]', first=True)
        email = r.html.xpath(f'//*[@id="main"]/div/div[3]/div[1]/div/div[{x}]/p[1]/strong[2]', first=True)
        
        clinics_list = {
            "Clinic": clinic.text,
            "Address": adress.text if adress else "None",
            "Phone": phone.text if phone else "None",
            "Email": email.text if email else "None"
        }
        all_clinics.append(clinics_list)
        print(clinics_list)

# 可选:将数据保存到JSON文件
# import json
# with open('clinics.json', 'w', encoding='utf-8') as f:
#     json.dump(all_clinics, f, ensure_ascii=False, indent=2)

关键优化点:

  • 自动识别总页数,避免无效循环
  • 每页遍历增加空值判断,提前终止不足30条数据的页面循环
  • 用字段存在性判断替代冗余的try-except

方案2:API请求提取完整数据

API返回的data字段包含地址、电话、邮箱等所有信息,直接从对应字段提取即可,以下是修改后的代码:

import requests

api_url = "https://is-api.dent.cz/api/v1/web/workplaces"
all_clinics = []

page = 1
while True:
    payload = {
        "deleted": False,
        "filter": "accepts_new_patients=false",
        "fulltext": "",
        "page": page,
        "per_page": 30,
        "sort_fields": "name",
    }
    
    response = requests.post(api_url, json=payload)
    data = response.json()
    
    if not data["data"]:  # 无数据则终止分页循环
        break
    
    for item in data["data"]:
        clinic_info = {
            "Clinic": item.get("name", "None"),
            "Address": f"{item.get('street', '')}, {item.get('city', '')}, {item.get('zip_code', '')}".strip(', '),
            "Phone": item.get("phone", "None"),
            "Email": item.get("email", "None"),
            "Website": item.get("web", "None")
        }
        all_clinics.append(clinic_info)
        print(clinic_info)
    
    page += 1

# 可选:将数据保存到JSON文件
# import json
# with open('clinics_api.json', 'w', encoding='utf-8') as f:
#     json.dump(all_clinics, f, ensure_ascii=False, indent=2)

关键优化点:

  • 用while True实现自动分页,直到无数据返回
  • 用item.get()方法避免KeyError,字段不存在时返回默认值
  • 拼接街道、城市、邮编为完整地址

内容的提问来源于stack exchange,提问作者Filip Chobodicky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 17:56:09