R中Hunspell无拼写建议时越界错误的解决方法
解决hunspell处理无拼写建议词汇时的越界错误问题
我尝试自动检查data.table/data.frame的字符串列拼写,发现当hunspell_suggest返回无建议的空列表(比如像“pippasnjfjsfiadjg”这种完全无意义的词汇)时,多种处理方法都会触发“越界”错误。我知道需要识别这些空建议,将其排除在“选取首个建议”的代码逻辑之外,但不知道具体怎么实现。
原测试代码如下:
library(dplyr) library(stringi) library(hunspell) df1 <- data.frame("Index" = 1:7, "Text" = c("pippasnjfjsfiadjg came to dinner with us tonigh.", "Wuld you like to trave with me?", "There is so muh to undestand.", "Sentences cone in many shaes and sizes.", "Learnin R is fun", "yesterday was Friday", "bing search engine"), stringsAsFactors = FALSE) # 获取拼写错误的词汇 badwords <- hunspell(df1$Text) %>% unlist # 提取每个错误词汇的首个拼写建议 suggestions <- sapply(hunspell_suggest(badwords), "[[", 1) # 替换错误词汇 mutate(df1, Text = stri_replace_all_fixed(str = Text, pattern = badwords, replacement = suggestions, vectorize_all = FALSE)) -> out
解决方案
核心是在提取拼写建议时,先判断每个建议列表是否为空:为空的情况直接保留原词(或返回NA,根据需求调整),避免直接取[[1]]触发越界错误。修改后的代码如下:
library(dplyr) library(stringi) library(hunspell) df1 <- data.frame("Index" = 1:7, "Text" = c("pippasnjfjsfiadjg came to dinner with us tonigh.", "Wuld you like to trave with me?", "There is so muh to undestand.", "Sentences cone in many shaes and sizes.", "Learnin R is fun", "yesterday was Friday", "bing search engine"), stringsAsFactors = FALSE) # 获取拼写错误的词汇 badwords <- hunspell(df1$Text) %>% unlist # 提取首个建议,空建议时返回原词 suggestions <- sapply(badwords, function(word) { sugg <- hunspell_suggest(word)[[1]] if(length(sugg) > 0) sugg[[1]] else word }) # 替换错误词汇 out <- mutate(df1, Text = stri_replace_all_fixed(str = Text, pattern = badwords, replacement = suggestions, vectorize_all = FALSE))
关键说明
- 替换了原代码中直接用
sapply(hunspell_suggest(badwords), "[[", 1)的逻辑,改为逐个处理每个错误词汇:先获取它的建议列表,若列表长度大于0则取首个建议,否则返回原词。 - 这样处理后,像“pippasnjfjsfiadjg”这类无拼写建议的词汇会被原样保留,不会触发越界错误,同时其他有建议的错误词汇会正常替换。
内容的提问来源于stack exchange,提问作者Magasinus
相关产品推荐
相关产品推荐

