You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用quanteda计算余弦相似度如何仅返回跨文档集配对结果

问题根源

你得到集合内配对结果的核心原因有两个:

  • 文件读入逻辑错误:新闻报道、政治决策文件的读入路径完全一致,两次readtext()调用都读取了目标目录下的全部文件,并未将两类文档拆分到独立语料对象中,两个dfm实际同时包含了新闻和决策两类文件。
  • dfm未做特征对齐:独立构建的两个dfm特征集合不完全匹配,会导致相似度计算时出现匹配偏差。
修复方案

可选读入方式(二选一即可)

  • 分文件夹存储(推荐):将两类文档分别存入两个独立子文件夹,新闻报道存放至txt_directory/news/,政治决策文件存放至txt_directory/policy/,读入时直接指定对应子文件夹路径。
  • 文件名规则过滤:如果不想调整文件夹结构,可以给两类文件设置统一命名前缀,例如新闻文件以news_开头、决策文件以pol_开头,读入时通过pattern参数过滤对应文件即可,无需移动文件。

修正后完整代码

library(quanteda)
library(readtext)
library(writexl)

# 方式1:分文件夹读入
news_articles <- readtext(paste0(txt_directory, "news/*"), encoding = "UTF-8")
pol_decisions <- readtext(paste0(txt_directory, "policy/*"), encoding = "UTF-8")

# 方式2:文件名规则过滤读入(替换上面两行读入代码即可)
# news_articles <- readtext(paste0(txt_directory, "*"), pattern = "^news_.*\\.txt$", encoding = "UTF-8")
# pol_decisions <- readtext(paste0(txt_directory, "*"), pattern = "^pol_.*\\.txt$", encoding = "UTF-8")

# 构建语料对象
news_corpus <- corpus(news_articles)
pol_corpus <- corpus(pol_decisions)

# 预处理新闻语料
news_toks <- tokens(
  news_corpus,
  what = "word",
  remove_punct = TRUE,
  remove_symbols = TRUE,
  remove_numbers = TRUE,
  remove_url = TRUE,
  remove_separators = TRUE,
  verbose = TRUE
) |> 
  tokens_tolower(keep_acronyms = FALSE) |> 
  tokens_select(stopwords("danish"), selection = "remove") |> 
  tokens_wordstem()

# 预处理政策语料
pol_toks <- tokens(
  pol_corpus,
  what = "word",
  remove_punct = TRUE,
  remove_symbols = TRUE,
  remove_numbers = TRUE,
  remove_url = TRUE,
  remove_separators = TRUE,
  verbose = TRUE
) |> 
  tokens_tolower(keep_acronyms = FALSE) |> 
  tokens_select(stopwords("danish"), selection = "remove") |> 
  tokens_wordstem()

# 构建dfm并对齐特征,保证两个dfm特征维度、顺序完全一致
news_dfm <- dfm(news_toks)
pol_dfm <- dfm(pol_toks)
pol_dfm_aligned <- dfm_match(pol_dfm, features = featnames(news_dfm))

# 计算跨集合余弦相似度:传入独立x和y参数时,仅返回x(新闻)与y(政策)的配对结果
cosine_res <- textstat_simil(
  news_dfm,
  y = pol_dfm_aligned,
  margin = "documents",
  method = "cosine"
)

# 结果排序导出
cosine_df <- as.data.frame(cosine_res)
cosine_df <- cosine_df[order(-cosine_df$cosine),]
write_xlsx(cosine_df, "新闻-政策相似度结果.xlsx")
说明

只要正确将两类文档拆分到独立的dfm对象中,textstat_simil()在同时接收到x和y两个独立dfm参数时,默认只会计算x内文档与y内文档的两两相似度,不会返回两个集合内部的配对结果,无需额外对结果做子集过滤。

内容的提问来源于stack exchange,提问作者Jonas Videbæk Jørgensen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 03:18:16