You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决R中stringi转义Unicode与easyPubMed十六进制编码匹配问题

解决方案

1. 消除R CMD检查的非ASCII警告

无需对大学列表的字符进行转义,直接在R包的DESCRIPTION文件中添加一行配置:

Encoding: UTF-8

R会自动识别包内数据为UTF-8编码,不再触发"found non-ASCII strings"警告,同时保留大学名称的原始带重音字符,为后续匹配提供便利。

2. 转换easyPubMed返回的HTML实体编码

easyPubMed返回的地址中é这类格式是HTML实体编码,不属于字符编码问题,因此修改articles_to_list的encoding参数无效。需将这些实体转换为对应的UTF-8字符,以下是两种可行方法:

方法一:用xml2包解析转换

xml2可自动解析HTML实体并转换为正常字符:

library(xml2)

# 定义转换函数
convert_html_entities <- function(text) {
  xml_text(read_html(paste0("<span>", text, "</span>")))
}

# 示例转换
raw_address <- "Laboratoire de Recherche sur le Comportement Social, D&#xe9;partement de Psychologie, Universit&#xe9; du Qu&#xe9;bec &#xe0; Montr&#xe9;al, Montr&#xe9;al, QC H3C 3P8, Canada."
clean_address <- convert_html_entities(raw_address)
clean_address
#> [1] "Laboratoire de Recherche sur le Comportement Social, Département de Psychologie, Université du Québec à Montréal, Montréal, QC H3C 3P8, Canada."

方法二:用stringi正则替换

通过正则匹配HTML实体的十六进制格式,手动转换为Unicode字符:

library(stringi)

convert_html_entities_stri <- function(text) {
  stri_replace_all_regex(
    text,
    "&#x([0-9a-fA-F]+);",
    function(match) {
      code <- as.integer(strtoi(match[2], base = 16))
      stri_unescape_unicode(sprintf("\\u%04x", code))
    },
    vectorize_all = FALSE
  )
}

# 示例转换
clean_address2 <- convert_html_entities_stri(raw_address)
clean_address2
#> [1] "Laboratoire de Recherche sur le Comportement Social, Département de Psychologie, Université du Québec à Montréal, Montréal, QC H3C 3P8, Canada."

3. 验证匹配

转换完成后,即可正常匹配带重音的字符:

# 假设大学列表中的名称
university <- "Université du Québec à Montréal"

# 匹配验证
grepl(university, clean_address)
#> [1] TRUE

内容的提问来源于stack exchange,提问作者rempsyc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 21:45:55