使用quanteda计算余弦相似度如何仅返回跨文档集配对结果
问题根源
你得到集合内配对结果的核心原因有两个:
- 文件读入逻辑错误:新闻报道、政治决策文件的读入路径完全一致,两次
readtext()调用都读取了目标目录下的全部文件,并未将两类文档拆分到独立语料对象中,两个dfm实际同时包含了新闻和决策两类文件。 - dfm未做特征对齐:独立构建的两个dfm特征集合不完全匹配,会导致相似度计算时出现匹配偏差。
修复方案
可选读入方式(二选一即可)
- 分文件夹存储(推荐):将两类文档分别存入两个独立子文件夹,新闻报道存放至
txt_directory/news/,政治决策文件存放至txt_directory/policy/,读入时直接指定对应子文件夹路径。 - 文件名规则过滤:如果不想调整文件夹结构,可以给两类文件设置统一命名前缀,例如新闻文件以
news_开头、决策文件以pol_开头,读入时通过pattern参数过滤对应文件即可,无需移动文件。
修正后完整代码
library(quanteda) library(readtext) library(writexl) # 方式1:分文件夹读入 news_articles <- readtext(paste0(txt_directory, "news/*"), encoding = "UTF-8") pol_decisions <- readtext(paste0(txt_directory, "policy/*"), encoding = "UTF-8") # 方式2:文件名规则过滤读入(替换上面两行读入代码即可) # news_articles <- readtext(paste0(txt_directory, "*"), pattern = "^news_.*\\.txt$", encoding = "UTF-8") # pol_decisions <- readtext(paste0(txt_directory, "*"), pattern = "^pol_.*\\.txt$", encoding = "UTF-8") # 构建语料对象 news_corpus <- corpus(news_articles) pol_corpus <- corpus(pol_decisions) # 预处理新闻语料 news_toks <- tokens( news_corpus, what = "word", remove_punct = TRUE, remove_symbols = TRUE, remove_numbers = TRUE, remove_url = TRUE, remove_separators = TRUE, verbose = TRUE ) |> tokens_tolower(keep_acronyms = FALSE) |> tokens_select(stopwords("danish"), selection = "remove") |> tokens_wordstem() # 预处理政策语料 pol_toks <- tokens( pol_corpus, what = "word", remove_punct = TRUE, remove_symbols = TRUE, remove_numbers = TRUE, remove_url = TRUE, remove_separators = TRUE, verbose = TRUE ) |> tokens_tolower(keep_acronyms = FALSE) |> tokens_select(stopwords("danish"), selection = "remove") |> tokens_wordstem() # 构建dfm并对齐特征,保证两个dfm特征维度、顺序完全一致 news_dfm <- dfm(news_toks) pol_dfm <- dfm(pol_toks) pol_dfm_aligned <- dfm_match(pol_dfm, features = featnames(news_dfm)) # 计算跨集合余弦相似度:传入独立x和y参数时,仅返回x(新闻)与y(政策)的配对结果 cosine_res <- textstat_simil( news_dfm, y = pol_dfm_aligned, margin = "documents", method = "cosine" ) # 结果排序导出 cosine_df <- as.data.frame(cosine_res) cosine_df <- cosine_df[order(-cosine_df$cosine),] write_xlsx(cosine_df, "新闻-政策相似度结果.xlsx")
说明
只要正确将两类文档拆分到独立的dfm对象中,textstat_simil()在同时接收到x和y两个独立dfm参数时,默认只会计算x内文档与y内文档的两两相似度,不会返回两个集合内部的配对结果,无需额外对结果做子集过滤。
内容的提问来源于stack exchange,提问作者Jonas Videbæk Jørgensen
相关产品推荐
相关产品推荐

