You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中提取application/ld+json内容时遇JSON解析垃圾后缀错误

解决Hemnet网站application/ld+json多片段JSON解析问题

我正在开展一项房地产经济学小型研究项目,需要从Hemnet房源页面的<script type="application/ld+json">标签中提取价格、地块面积、描述、位置等数据。因为Hemnet聚合了多网站房源,不同来源的CSS选择器差异大,所以选择通过结构化的LD-JSON获取数据,方便后续复用。

我尝试了以下代码:

library(rvest)
library(xml2)
library(jsonlite)
library(dplyr)

o_url <- "https://www.hemnet.se/bostad/tomt-lisselbo-falu-kommun-svartskar-1-17-14704536"
o_html <- read_html(o_url)

o_json <- html_nodes(o_html, "[type=\"application/ld+json\"]") %>% html_text()
ldjson <- jsonlite::fromJSON(o_json)

运行后出现解析错误:

Error: parse error: trailing garbage
          https://www.hemnet.se"   }   {     "@context": "http://schem
                     (right here) ------^

问题原因

你提取的o_json是多个独立的JSON对象拼接在一起的文本,而jsonlite::fromJSON()只能解析单个合法的JSON对象或数组。直接合并这些片段会违反JSON语法规则,导致解析失败。

解决方案

需要将每个<script type="application/ld+json">节点的文本单独提取并解析,再筛选出包含房源数据的对象:

library(rvest)
library(jsonlite)
library(stringr)
library(purrr)

# 目标房源URL
o_url <- "https://www.hemnet.se/bostad/tomt-lisselbo-falu-kommun-svartskar-1-17-14704536"
o_html <- read_html(o_url)

# 提取所有LD-JSON节点的文本内容(每个节点对应一个独立JSON片段)
json_fragments <- html_nodes(o_html, "[type=\"application/ld+json\"]") %>% 
  html_text()

# 逐个解析每个JSON片段,跳过解析失败的无效内容
parsed_results <- map(json_fragments, function(fragment) {
  clean_fragment <- str_squish(fragment) # 清理多余空格和换行
  tryCatch(fromJSON(clean_fragment), error = function(e) NULL)
}) %>% 
  discard(is.null) # 移除解析失败的空对象

# 筛选出房源相关的结构化数据(通常标记为RealEstateListing类型)
property_data <- parsed_results %>% 
  keep(function(x) x[["@type"]] == "RealEstateListing") %>% 
  pluck(1) # 取第一个匹配的房源数据

# 提取需要的核心字段(字段名可根据实际返回结果调整)
extracted_info <- list(
  价格 = property_data$price,
  地块面积 = property_data$plotSize$value,
  房源描述 = property_data$description,
  位置 = property_data$address$addressLocality
)

print(extracted_info)

代码优化建议

  • 容错机制:加入tryCatch确保单个无效JSON片段不会导致整个脚本中断
  • 精准筛选:通过@type字段定位房源数据,避免混入网站自身的结构化信息(比如站点导航、商家信息等)
  • 字段适配:Hemnet不同房源的LD-JSON字段可能略有差异,可根据实际返回结果调整plotSize/floorSize等字段名
  • 批量扩展:如果需要爬取多个房源,可将上述逻辑封装成函数,传入URL列表实现批量数据提取

内容的提问来源于stack exchange,提问作者chappe29

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 23:27:35