You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言爬取URL无变化动态分页医生名录的问题咨询

爬取实现方案

该站点为前端动态渲染的单页应用,分页时通过后台接口异步加载数据,URL不会更新,直接调用后台分页接口即可高效获取全量结构化数据,完整实现代码如下:

library(dplyr)
library(rvest)
library(httr)
library(stringr)

# 安装依赖包命令:install.packages(c("dplyr","rvest","httr","stringr"))

total_page <- 6
all_doctors_df <- data.frame()

for (current_page in 1:total_page) {
  # 构造分页POST请求
  resp <- POST(
    url = "https://www.doctorisy.com/ServerSide/functionSearch.asmx/GetFilterDoctors",
    content_type("application/json; charset=utf-8"),
    body = list(
      country = "guatemala",
      state = "",
      city = "",
      speciality = "",
      insurance = "",
      name = "",
      page = current_page,
      itemsPerPage = 12
    ),
    encode = "json"
  )
  
  # 解析返回的HTML片段
  page_html <- content(resp)$d %>% 
    read_html()
  
  # 定位每个医生的独立卡片节点,解决原代码数据错位问题
  doctor_cards <- page_html %>% 
    html_nodes(".card-doctor.mat-card")
  
  # 提取结构化字段
  page_df <- lapply(doctor_cards, function(card) {
    data.frame(
      doctor_name = html_node(card, ".title-primary") %>% html_text(trim = TRUE),
      speciality = html_node(card, ".speciality") %>% html_text(trim = TRUE),
      clinic_name = html_node(card, ".clinic-name") %>% html_text(trim = TRUE),
      address = html_node(card, ".address") %>% html_text(trim = TRUE),
      phone = html_node(card, ".phone a") %>% html_attr("href") %>% str_remove("tel:"),
      stringsAsFactors = FALSE
    )
  }) %>% bind_rows()
  
  all_doctors_df <- bind_rows(all_doctors_df, page_df)
  # 控制请求频率避免被拦截
  Sys.sleep(1.5)
}

# 查看最终全量结构化数据,每一行对应1条医生信息
View(all_doctors_df)

方案说明

  • 直接调用站点后台分页接口,无需模拟浏览器渲染,爬取效率更高
  • 选择器精准定位单个医生卡片节点,解决了原代码的数据错位问题
  • 直接提取各字段结构化内容,无需处理非结构化全文本,输出结果为标准数据框格式,可直接导出为csv等格式使用

内容的提问来源于stack exchange,提问作者Duck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 05:36:01