运行DocumentTermMatrix(DTM)时数据空格消失问题求助
问题与解决方法
问题详情
- 已完成的预处理操作:移除文本中的标点与数字、将内容统一转为小写、将数据集从xls格式转换为csv格式
- 运行
DocumentTermMatrix后得到的异常结果:
adelaidetolondonviasingaporewasmyfirstexperienceofbusinessclassonsingaporeairlinesaandiwasdisappointedfoodifecrewetcwerefinethoughnothingexceptionalbuttheseating
- 原始数据集内容:
Adelaide to London via Singapore was my first experience of Business Class on Singapore Airlines A380, and I was disappointed. Food, IFE, crew etc. were fine (though nothing exceptional) but the seating
核心原因与修复步骤
空格消失的问题大概率是预处理环节误删了空格,或是构建语料库时未正确保留文本的分词结构,按以下步骤排查修复:
修正标点/数字移除的正则表达式
如果你用gsub做文本清理,别使用过于宽泛的正则规则(比如gsub("[^a-z]","", text)会把空格一并删除)。正确写法需保留空格,示例代码:# 仅移除标点、数字,保留字母和空格 cleaned_text <- gsub("[^a-z\\s]", "", tolower(raw_text))规范语料库构建流程
使用tm包构建语料库时,确保文本以正确的字符串形式传入,避免内容被错误合并。示例流程:library(tm) # 读取csv数据,假设文本列名为text df <- read.csv("your_data.csv", stringsAsFactors = FALSE) # 构建语料库 corpus <- VCorpus(VectorSource(df$text)) # 应用正确的预处理逻辑 corpus <- tm_map(corpus, content_transformer(function(x) { gsub("[^a-z\\s]", "", tolower(x)) })) # 生成DocumentTermMatrix dtm <- DocumentTermMatrix(corpus)验证每一步的中间结果
在预处理完成后直接打印清理后的文本,确认空格是否保留。提前在中间环节排查问题,避免等生成DTM后再追溯错误。
内容的提问来源于stack exchange,提问作者Nicholas Kim
相关产品推荐
相关产品推荐

