You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行DocumentTermMatrix(DTM)时数据空格消失问题求助

问题与解决方法

问题详情

  • 已完成的预处理操作:移除文本中的标点与数字、将内容统一转为小写、将数据集从xls格式转换为csv格式
  • 运行DocumentTermMatrix后得到的异常结果:

adelaidetolondonviasingaporewasmyfirstexperienceofbusinessclassonsingaporeairlinesaandiwasdisappointedfoodifecrewetcwerefinethoughnothingexceptionalbuttheseating

  • 原始数据集内容:

Adelaide to London via Singapore was my first experience of Business Class on Singapore Airlines A380, and I was disappointed. Food, IFE, crew etc. were fine (though nothing exceptional) but the seating

核心原因与修复步骤

空格消失的问题大概率是预处理环节误删了空格,或是构建语料库时未正确保留文本的分词结构,按以下步骤排查修复:

  1. 修正标点/数字移除的正则表达式
    如果你用gsub做文本清理,别使用过于宽泛的正则规则(比如gsub("[^a-z]","", text)会把空格一并删除)。正确写法需保留空格,示例代码:

    # 仅移除标点、数字,保留字母和空格
    cleaned_text <- gsub("[^a-z\\s]", "", tolower(raw_text))
    
  2. 规范语料库构建流程
    使用tm包构建语料库时,确保文本以正确的字符串形式传入,避免内容被错误合并。示例流程:

    library(tm)
    # 读取csv数据,假设文本列名为text
    df <- read.csv("your_data.csv", stringsAsFactors = FALSE)
    # 构建语料库
    corpus <- VCorpus(VectorSource(df$text))
    # 应用正确的预处理逻辑
    corpus <- tm_map(corpus, content_transformer(function(x) {
      gsub("[^a-z\\s]", "", tolower(x))
    }))
    # 生成DocumentTermMatrix
    dtm <- DocumentTermMatrix(corpus)
    
  3. 验证每一步的中间结果
    在预处理完成后直接打印清理后的文本,确认空格是否保留。提前在中间环节排查问题,避免等生成DTM后再追溯错误。


内容的提问来源于stack exchange,提问作者Nicholas Kim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:55:21