在R中提取application/ld+json内容时遇JSON解析垃圾后缀错误
解决Hemnet网站application/ld+json多片段JSON解析问题
我正在开展一项房地产经济学小型研究项目,需要从Hemnet房源页面的<script type="application/ld+json">标签中提取价格、地块面积、描述、位置等数据。因为Hemnet聚合了多网站房源,不同来源的CSS选择器差异大,所以选择通过结构化的LD-JSON获取数据,方便后续复用。
我尝试了以下代码:
library(rvest) library(xml2) library(jsonlite) library(dplyr) o_url <- "https://www.hemnet.se/bostad/tomt-lisselbo-falu-kommun-svartskar-1-17-14704536" o_html <- read_html(o_url) o_json <- html_nodes(o_html, "[type=\"application/ld+json\"]") %>% html_text() ldjson <- jsonlite::fromJSON(o_json)
运行后出现解析错误:
Error: parse error: trailing garbage https://www.hemnet.se" } { "@context": "http://schem (right here) ------^
问题原因
你提取的o_json是多个独立的JSON对象拼接在一起的文本,而jsonlite::fromJSON()只能解析单个合法的JSON对象或数组。直接合并这些片段会违反JSON语法规则,导致解析失败。
解决方案
需要将每个<script type="application/ld+json">节点的文本单独提取并解析,再筛选出包含房源数据的对象:
library(rvest) library(jsonlite) library(stringr) library(purrr) # 目标房源URL o_url <- "https://www.hemnet.se/bostad/tomt-lisselbo-falu-kommun-svartskar-1-17-14704536" o_html <- read_html(o_url) # 提取所有LD-JSON节点的文本内容(每个节点对应一个独立JSON片段) json_fragments <- html_nodes(o_html, "[type=\"application/ld+json\"]") %>% html_text() # 逐个解析每个JSON片段,跳过解析失败的无效内容 parsed_results <- map(json_fragments, function(fragment) { clean_fragment <- str_squish(fragment) # 清理多余空格和换行 tryCatch(fromJSON(clean_fragment), error = function(e) NULL) }) %>% discard(is.null) # 移除解析失败的空对象 # 筛选出房源相关的结构化数据(通常标记为RealEstateListing类型) property_data <- parsed_results %>% keep(function(x) x[["@type"]] == "RealEstateListing") %>% pluck(1) # 取第一个匹配的房源数据 # 提取需要的核心字段(字段名可根据实际返回结果调整) extracted_info <- list( 价格 = property_data$price, 地块面积 = property_data$plotSize$value, 房源描述 = property_data$description, 位置 = property_data$address$addressLocality ) print(extracted_info)
代码优化建议
- 容错机制:加入
tryCatch确保单个无效JSON片段不会导致整个脚本中断 - 精准筛选:通过
@type字段定位房源数据,避免混入网站自身的结构化信息(比如站点导航、商家信息等) - 字段适配:Hemnet不同房源的LD-JSON字段可能略有差异,可根据实际返回结果调整
plotSize/floorSize等字段名 - 批量扩展:如果需要爬取多个房源,可将上述逻辑封装成函数,传入URL列表实现批量数据提取
内容的提问来源于stack exchange,提问作者chappe29
相关产品推荐
相关产品推荐

