You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中保留特殊字符构建DocumentTermMatrix,区分LKC与LKC_?

Got it, this is a common issue when working with text mining in R—default tokenization often strips out underscores, which makes terms like LKC and LKC_ get merged. Let's walk through how to fix this with the two most popular text mining packages: tm and quanteda.


Using the tm Package

The default tm setup treats underscores as word separators or punctuation, so we need to tweak both punctuation handling and tokenization to keep them intact.

  1. Load the package and create sample data
library(tm)

# Sample texts containing your target distinct terms
sample_texts <- c("I referenced LKC in the report", "LKC_ is a separate identifier")
my_corpus <- VCorpus(VectorSource(sample_texts))
  1. Customize punctuation removal (keep underscores)
    The default removePunctuation function strips underscores, so we'll write a custom function to only remove other punctuation:
# Function to remove all punctuation EXCEPT underscores
remove_punct_keep_underscore <- function(x) {
  gsub("[^[:alnum:][:space:]_]", "", x)
}

# Apply the custom function to the corpus
my_corpus <- tm_map(my_corpus, content_transformer(remove_punct_keep_underscore))
  1. Use a custom tokenizer (split only on whitespace)
    The built-in tokenizer splits words at underscores, so we'll use a simple whitespace-based tokenizer instead:
# Tokenizer that splits text only on spaces, preserving underscores
custom_tokenizer <- function(x) {
  unlist(strsplit(as.character(x), "\\s+"))
}
  1. Build and verify the DocumentTermMatrix
# Construct DTM with our custom settings
my_dtm <- DocumentTermMatrix(my_corpus, control = list(tokenize = custom_tokenizer))

# Inspect the result - you'll see both `LKC` and `LKC_` as separate columns
inspect(my_dtm)

Using the quanteda Package

Quanteda's default tokenization is more flexible and preserves underscores out of the box, making this the simpler option for most cases.

  1. Load the package and process text
library(quanteda)

# Sample texts
sample_texts <- c("I referenced LKC in the report", "LKC_ is a separate identifier")

# Tokenize text (default settings retain underscores)
text_tokens <- tokens(sample_texts)

# Build a Document-Feature Matrix (quanteda's equivalent of DTM)
my_dfm <- dfm(text_tokens)

# Check the features - `LKC` and `LKC_` will appear as distinct entries
print(my_dfm)

If you need to adjust for other special characters, you can tweak the tokens() function parameters (e.g., remove_punct = FALSE to keep all punctuation) or use tokens_remove() to target specific characters.


内容的提问来源于stack exchange,提问作者jz_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:25:13