R语言中Quanteda创建Corpus与DFM后文档数量减少的问题求助
Hey there, let's walk through the most likely reasons your corpus or DFM is missing documents, plus how to fix each issue:
1. Empty or Whitespace-Only Files Are Being Removed
Quanteda automatically filters out documents that have no content (or only spaces/newlines) by default when creating a corpus. This is a sensible default, but it might catch you off guard if you have files that look "empty" but you intended to keep them.
Fixes:
- When creating your corpus, disable the empty-document filter with the
remove_emptyparameter:my_corpus <- corpus(my_readtext_data, remove_empty = FALSE) - After creating the corpus, identify which docs are empty to decide whether to keep or remove them:
# Find docs with zero tokens empty_docs <- which(ntoken(my_corpus) == 0) # View their filenames (if using readtext) docvars(my_corpus, "doc_id")[empty_docs]
2. Files Failed to Read Properly
If some files have encoding issues, are corrupted, or are in a format Quanteda can't parse, they might get skipped during the read process (often with a warning you might have missed!).
Fixes:
- First, verify how many files you actually read vs. expected:
# Check number of files in your directory length(list.files("your_file_directory", pattern = "*.txt")) # adjust pattern as needed # Check number of docs in your readtext object my_readtext <- readtext("your_file_directory/*.txt") nrow(my_readtext) - If counts don't match, specify the correct file encoding (common issues with non-UTF-8 files):
my_readtext <- readtext("your_file_directory/*.txt", encoding = "Latin-1") # or "UTF-16", etc. - Keep an eye on the R console for warnings about unreadable files—they'll tell you exactly which ones failed.
3. DFM Construction Filtered Out Token-Less Documents
Even if your corpus has all documents, when building a DFM, any document that ends up with zero tokens (after tokenization, removing stopwords, trimming rare terms, etc.) will be dropped by default.
Fixes:
- Keep these empty rows in your DFM by setting
remove_empty = FALSEindfm():my_dfm <- dfm(my_corpus, remove_empty = FALSE) - Pre-process your corpus to avoid empty docs first:
# Remove stopwords but avoid stripping all content my_corpus_clean <- tokens(my_corpus) %>% tokens_remove(stopwords("en")) %>% tokens_remove(pattern = "^\\p{Punct}+$", valuetype = "regex") # keep meaningful terms # Check token counts again before building DFM ntoken(my_corpus_clean)
4. File Path/Pattern Matching Mistakes
It's easy to accidentally mismatch the files you're trying to read—maybe your wildcard pattern is wrong, or you're not including subdirectories.
Fixes:
- Double-check your file path and pattern:
# List all files matching your pattern to confirm list.files("your_file_directory", pattern = "*.txt", full.names = TRUE) - If you need to include files in subfolders, use
recursive = TRUE:my_readtext <- readtext("your_file_directory/**/*.txt", recursive = TRUE)
内容的提问来源于stack exchange,提问作者Nicholas Bradley

