如何在R中保留特殊字符构建DocumentTermMatrix,区分LKC与LKC_?
Got it, this is a common issue when working with text mining in R—default tokenization often strips out underscores, which makes terms like LKC and LKC_ get merged. Let's walk through how to fix this with the two most popular text mining packages: tm and quanteda.
tm Package The default tm setup treats underscores as word separators or punctuation, so we need to tweak both punctuation handling and tokenization to keep them intact.
- Load the package and create sample data
library(tm) # Sample texts containing your target distinct terms sample_texts <- c("I referenced LKC in the report", "LKC_ is a separate identifier") my_corpus <- VCorpus(VectorSource(sample_texts))
- Customize punctuation removal (keep underscores)
The defaultremovePunctuationfunction strips underscores, so we'll write a custom function to only remove other punctuation:
# Function to remove all punctuation EXCEPT underscores remove_punct_keep_underscore <- function(x) { gsub("[^[:alnum:][:space:]_]", "", x) } # Apply the custom function to the corpus my_corpus <- tm_map(my_corpus, content_transformer(remove_punct_keep_underscore))
- Use a custom tokenizer (split only on whitespace)
The built-in tokenizer splits words at underscores, so we'll use a simple whitespace-based tokenizer instead:
# Tokenizer that splits text only on spaces, preserving underscores custom_tokenizer <- function(x) { unlist(strsplit(as.character(x), "\\s+")) }
- Build and verify the DocumentTermMatrix
# Construct DTM with our custom settings my_dtm <- DocumentTermMatrix(my_corpus, control = list(tokenize = custom_tokenizer)) # Inspect the result - you'll see both `LKC` and `LKC_` as separate columns inspect(my_dtm)
quanteda Package Quanteda's default tokenization is more flexible and preserves underscores out of the box, making this the simpler option for most cases.
- Load the package and process text
library(quanteda) # Sample texts sample_texts <- c("I referenced LKC in the report", "LKC_ is a separate identifier") # Tokenize text (default settings retain underscores) text_tokens <- tokens(sample_texts) # Build a Document-Feature Matrix (quanteda's equivalent of DTM) my_dfm <- dfm(text_tokens) # Check the features - `LKC` and `LKC_` will appear as distinct entries print(my_dfm)
If you need to adjust for other special characters, you can tweak the tokens() function parameters (e.g., remove_punct = FALSE to keep all punctuation) or use tokens_remove() to target specific characters.
内容的提问来源于stack exchange,提问作者jz_

