You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用R爬取指定网页数据并匹配字段后输出为data frame

解决思路

你之前的问题是单独全局抓取三类元素,没有考虑「每个房型对应多个餐食/价格组合」的层级关系,导致长度不匹配、无法对齐。正确的逻辑是先拆分到每个独立的房型节点,再逐个节点内提取对应的餐食和价格,就能保证对应关系完全匹配网页展示逻辑。

完整可运行代码

# 加载所需包
library(rvest)
library(dplyr)
library(stringr)
library(purrr)

# 仅读取1次网页,避免重复请求和数据不一致
target_url <- "https://www.hotelissima.fr/s/h/ile-maurice/mahebourg/astroea-beach.html?searchType=accomodation&searchId=4&guideId=&filters=&withFlights=false&airportCode=PAR&airport=Paris&search=astroea%20beach&startdate=08%2F11%2F2021&stopdate=15%2F11%2F2021&duration=7&travelers=En%20couple&travelType=&rooms%5B0%5D.nbAdults=2&rooms%5B0%5D.nbChilds=0&rooms%5B0%5D.birthdates%5B0%5D=&rooms%5B0%5D.birthdates%5B1%5D=&rooms%5B0%5D.birthdates%5B2%5D=&rooms%5B0%5D.birthdates%5B3%5D=&rooms%5B0%5D.birthdates%5B4%5D="
page_html <- read_html(target_url)

# 提取所有房型节点,逐个节点内提取对应信息
result <- page_html %>%
  html_elements(".room") %>%
  map_dfr(function(room_node){
    # 提取当前房型名称
    room_type <- room_node %>% 
      html_element("h3") %>% 
      html_text() %>% 
      str_squish()
    # 提取当前房型下的所有餐食方案
    meal_plans <- room_node %>% 
      html_elements(".meal-plan-title") %>% 
      html_text() %>% 
      str_squish()
    # 提取当前房型下的所有价格
    prices <- room_node %>% 
      html_elements(".price") %>% 
      html_text() %>% 
      str_squish()
    # 拼成当前房型的子数据框
    tibble(
      RoomType = room_type,
      MealPlan = meal_plans,
      Price = prices
    )
  })

# 查看最终结果
print(result)

运行后得到的result就是你要求格式的data frame,完全匹配网页展示的对应关系。

内容的提问来源于stack exchange,提问作者user3115933

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 18:27:06