You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中如何基于文件语言标记排除指定数据文件?

Hey there! As a law student new to R, I totally get how overwhelming this can feel at first—no need to apologize for any rough wording 😊 Let’s break down how to exclude those French documents marked with "FR" efficiently, either before or after importing your files.

解决方法:排除含"FR"的法语文档

方法1:导入前筛选文件(更高效)

This is the better approach because it avoids importing unnecessary files entirely, saving memory and processing time—perfect for 1500+ documents. We’ll filter out files with "FR" in their names first, then import only the valid ones:

library(readtext)

folder <- "C:/[pathway]"
# Get all .txt files in the folder, then exclude those with "FR" in the filename
valid_files <- list.files(path = folder, pattern = "\\.txt$", full.names = TRUE)
valid_files <- valid_files[!grepl("FR", valid_files)]

# Import only the filtered files
submissions <- readtext(valid_files)

方法2:导入后筛选数据

If you’ve already imported all files, you can still filter out the French records easily. There are two straightforward ways to do this:

方式A:Filter based on doc_id

Use the original filename stored in doc_id to exclude entries with "FR":

# Remove rows where doc_id contains "FR"
submissions_filtered <- submissions[!grepl("FR", submissions$doc_id), ]

方式B:Filter based on your extracted language field

Since you’re already parsing the language from the filename, you can use that field directly after running your existing code:

# First run your original loop to extract submission_language, then filter
submissions_filtered <- submissions[submissions$submission_language != "FR", ]

Bonus: Optimize your existing code (replace the loop!)

Loops can be slow with 1500 entries. Using tidyr and dplyr will make your code cleaner and faster, plus it auto-handles data types:

library(readtext)
library(dplyr)
library(tidyr)

folder <- "C:/[pathway]"
submissions <- readtext(folder)

# Split doc_id into your desired columns in one step
submissions <- submissions %>%
  mutate(doc_id = gsub("\\.txt$", "", doc_id)) %>%  # Remove .txt suffix
  separate(doc_id, 
           into = c("submission_number", "submission_person", "submission_code", "submission_language", "submission_location"),
           sep = "_|-",  # Split on either _ or -
           convert = TRUE) %>%  # Auto-convert submission_number to numeric
  arrange(submission_number)  # Sort by submission number

# Now filter out French documents with one line
submissions_filtered <- submissions %>%
  filter(submission_language != "FR")

This version eliminates the manual loop and makes your workflow more reproducible for future analyses (like your upcoming sentiment analysis!).

内容的提问来源于stack exchange,提问作者televised-god

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:03:39