在R中如何基于文件语言标记排除指定数据文件?
Hey there! As a law student new to R, I totally get how overwhelming this can feel at first—no need to apologize for any rough wording 😊 Let’s break down how to exclude those French documents marked with "FR" efficiently, either before or after importing your files.
方法1:导入前筛选文件(更高效)
This is the better approach because it avoids importing unnecessary files entirely, saving memory and processing time—perfect for 1500+ documents. We’ll filter out files with "FR" in their names first, then import only the valid ones:
library(readtext) folder <- "C:/[pathway]" # Get all .txt files in the folder, then exclude those with "FR" in the filename valid_files <- list.files(path = folder, pattern = "\\.txt$", full.names = TRUE) valid_files <- valid_files[!grepl("FR", valid_files)] # Import only the filtered files submissions <- readtext(valid_files)
方法2:导入后筛选数据
If you’ve already imported all files, you can still filter out the French records easily. There are two straightforward ways to do this:
方式A:Filter based on doc_id
Use the original filename stored in doc_id to exclude entries with "FR":
# Remove rows where doc_id contains "FR" submissions_filtered <- submissions[!grepl("FR", submissions$doc_id), ]
方式B:Filter based on your extracted language field
Since you’re already parsing the language from the filename, you can use that field directly after running your existing code:
# First run your original loop to extract submission_language, then filter submissions_filtered <- submissions[submissions$submission_language != "FR", ]
Bonus: Optimize your existing code (replace the loop!)
Loops can be slow with 1500 entries. Using tidyr and dplyr will make your code cleaner and faster, plus it auto-handles data types:
library(readtext) library(dplyr) library(tidyr) folder <- "C:/[pathway]" submissions <- readtext(folder) # Split doc_id into your desired columns in one step submissions <- submissions %>% mutate(doc_id = gsub("\\.txt$", "", doc_id)) %>% # Remove .txt suffix separate(doc_id, into = c("submission_number", "submission_person", "submission_code", "submission_language", "submission_location"), sep = "_|-", # Split on either _ or - convert = TRUE) %>% # Auto-convert submission_number to numeric arrange(submission_number) # Sort by submission number # Now filter out French documents with one line submissions_filtered <- submissions %>% filter(submission_language != "FR")
This version eliminates the manual loop and makes your workflow more reproducible for future analyses (like your upcoming sentiment analysis!).
内容的提问来源于stack exchange,提问作者televised-god

