解决R中stringi转义Unicode与easyPubMed十六进制编码匹配问题
解决方案
1. 消除R CMD检查的非ASCII警告
无需对大学列表的字符进行转义,直接在R包的DESCRIPTION文件中添加一行配置:
Encoding: UTF-8
R会自动识别包内数据为UTF-8编码,不再触发"found non-ASCII strings"警告,同时保留大学名称的原始带重音字符,为后续匹配提供便利。
2. 转换easyPubMed返回的HTML实体编码
easyPubMed返回的地址中é这类格式是HTML实体编码,不属于字符编码问题,因此修改articles_to_list的encoding参数无效。需将这些实体转换为对应的UTF-8字符,以下是两种可行方法:
方法一:用xml2包解析转换
xml2可自动解析HTML实体并转换为正常字符:
library(xml2) # 定义转换函数 convert_html_entities <- function(text) { xml_text(read_html(paste0("<span>", text, "</span>"))) } # 示例转换 raw_address <- "Laboratoire de Recherche sur le Comportement Social, Département de Psychologie, Université du Québec à Montréal, Montréal, QC H3C 3P8, Canada." clean_address <- convert_html_entities(raw_address) clean_address #> [1] "Laboratoire de Recherche sur le Comportement Social, Département de Psychologie, Université du Québec à Montréal, Montréal, QC H3C 3P8, Canada."
方法二:用stringi正则替换
通过正则匹配HTML实体的十六进制格式,手动转换为Unicode字符:
library(stringi) convert_html_entities_stri <- function(text) { stri_replace_all_regex( text, "&#x([0-9a-fA-F]+);", function(match) { code <- as.integer(strtoi(match[2], base = 16)) stri_unescape_unicode(sprintf("\\u%04x", code)) }, vectorize_all = FALSE ) } # 示例转换 clean_address2 <- convert_html_entities_stri(raw_address) clean_address2 #> [1] "Laboratoire de Recherche sur le Comportement Social, Département de Psychologie, Université du Québec à Montréal, Montréal, QC H3C 3P8, Canada."
3. 验证匹配
转换完成后,即可正常匹配带重音的字符:
# 假设大学列表中的名称 university <- "Université du Québec à Montréal" # 匹配验证 grepl(university, clean_address) #> [1] TRUE
内容的提问来源于stack exchange,提问作者rempsyc
相关产品推荐
相关产品推荐

