You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Quanteda读取多文本文件遇阻:需清理#标记文本并解决向量报错

问题:批量导入文本到Quanteda并清理#标记内容

数据下载说明

尝试用readtext下载目标文本文件时,收到「远程URL无已知扩展名,请手动下载」的错误,因此使用rvest编写了批量下载20个文件的代码:

suppressPackageStartupMessages({
library(rvest)
})

# destination directory, change this at will
dest_dir <- "~/Temp"

# first get the two subfolders from the Data webpage
link <- "http://home.brisnet.org.au/~bgreen/Data/"
page <- read_html(link)
page %>%
html_elements("a") %>%
html_text() %>%
grep("/$", ., value = TRUE) -> sub_folder

# create relevant disk sub-directories, if
# they do not exist yet
for(subf in sub_folder) {
d <- file.path(dest_dir, subf)
if(!dir.exists(d)) {
success <- dir.create(d)
msg <- paste("created directory", d, "-", success)
message(msg)
  }
}

# prepare to download the files
dest_dir <- file.path(dest_dir, sub_folder)
source_url <- paste0(link, sub_folder)

success <- mapply(\(src, dest) {
# read each Data subfolder
# and get the file names therein
# then lapply 'download.file' to each filename
pg <- read_html(src)
pg %>%
html_elements("a") %>%
html_text() %>%
grep("\\.txt$", ., value = TRUE) %>%
lapply(\(x) {
s <- paste0(src, x)
d <- file.path(dest, x)
tryCatch(
download.file(url = s, destfile = d),
warning = function(w) w,
error = function(e) e
)
})
}, source_url, dest_dir)

lengths(success)  

文本清理与导入问题

需要先移除文本中#标记之间的无用文本,编写了以下stringi代码,但不确定是否能清理所有目标内容:

library("stringi")
toks <- stringi::stri_replace_all_regex(x, "#.*#\n{2}", "") |> tokens()

此外,将清理后的文本导入Quanteda时,遇到报错:「参数并非原子向量;正在强制转换」

求助需求

  • 验证上述stringi正则表达式是否能正确移除所有#标记之间的文本
  • 解决Quanteda导入时的「参数并非原子向量;正在强制转换」报错
  • 寻求更高效的多文本文件批量处理/复现方法

内容的提问来源于stack exchange,提问作者bgreen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 10:13:21