You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

文本相似度计算与去重方案咨询:余弦距离是否最优?

回答

Great question! Let's start with a direct answer: Using cosine distance (paired with text vectorization) is an extremely effective approach for identifying these near-duplicate texts, though "optimal" depends on your specific use case. For your scenario—where duplicates have minor vocabulary differences or partial content omissions—it’s absolutely a top-tier solution, and one of the most commonly used methods for this kind of problem.

Why unique() falls short

unique(text) only catches exact matches, but your data has near-duplicates (like the two football entries, where one is just missing the final sentence). These require a similarity-based approach, and cosine distance excels at capturing the "highly similar but not identical" relationship by measuring overlap between text vectors.

Step-by-step R implementation

For your dataset, we'll use TF-IDF to convert text into numerical vectors, calculate pairwise cosine similarity, filter near-duplicates with a threshold, and keep the longest text in each duplicate group. Here's the full workflow:

1. Load required packages

library(tm)
library(proxy)
library(dplyr)

2. Prepare your dataset & preprocess text

# Your original dataset
text <- c("Football is a family of team sports that involve, to varying degrees, kicking a ball to score a goal. Unqualified, the word football is understood to refer to whichever form of football is the most popular in the regional context in which the word appears. Sports commonly called football in certain places include association football (known as soccer in some countries); gridiron football (specifically American football or Canadian football); Australian rules football; rugby football (either rugby league or rugby union); and Gaelic football.[1][2] These different variations of football are known as football codes.", 
          "Football is a family of team sports that involve, to varying degrees, kicking a ball to score a goal. Unqualified, the word football is understood to refer to whichever form of football is the most popular in the regional context in which the word appears. Sports commonly called football in certain places include association football (known as soccer in some countries); gridiron football (specifically American football or Canadian football); Australian rules football; rugby football (either rugby league or rugby union); and Gaelic football.[1][2]", 
          "Tennis is a racket sport that can be played individually against a single opponent (singles) or between two teams of two players each (doubles). Each player uses a tennis racket that is strung with cord to strike a hollow rubber ball covered with felt over or around a net and into the opponent's court. The object of the game is to maneuver the ball in such a way that the opponent is not able to play a valid return. The player who is unable to return the ball will not gain a point, while the opposite player will.", 
          "Tennis is a racket sport that can be played individually against a single opponent (singles) or between two teams of two players each (doubles). Each player uses a tennis racket that is strung with cord to strike a hollow rubber ball covered with felt over or around a net and into the opponent's court. The object of the game is to maneuver the ball in such a way that the opponent is not able to play a valid return.", 
          "Rugby refers to the team sports rugby league and rugby union. Legend claims that rugby football was started about 1845 in Rugby School, Rugby, Warwickshire, England, although forms of football in which the ball was carried and tossed date to medieval times. Rugby eventually split into two sports in 1895 when twenty-one clubs split from the original Rugby Football Union, to form the Northern Union (later to be named rugby league in 1922) in the George Hotel, Huddersfield, Northern England over the issue of payment to players, thus making rugby league the first code to turn professional and pay its players, rugby union turned fully professional in 1995. Both sports are run by their respective world governing bodies World Rugby (rugby union) and the Rugby League International Federation (rugby league). Rugby football was one of many versions of football played at English public schools in the 19th century.[1][2] Although rugby league initially used rugby union rules, they are now wholly separate sports. In addition to these two codes, both American and Canadian football evolved from rugby football.")
ID <- c("Foot123", "Foot123", "Ten123", "Ten123", "Rugby123")
data <- data.frame(text, ID, stringsAsFactors = FALSE)

# Clean text: lowercase, remove punctuation/stopwords, trim whitespace
corpus <- VCorpus(VectorSource(data$text)) %>%
  tm_map(content_transformer(tolower)) %>%
  tm_map(removePunctuation) %>%
  tm_map(removeWords, stopwords("english")) %>%
  tm_map(stripWhitespace)

# Generate TF-IDF matrix (weights words by importance)
tfidf <- DocumentTermMatrix(corpus, control = list(weighting = weightTfIdf))
tfidf_mat <- as.matrix(tfidf)

3. Calculate cosine similarity & deduplicate

# Compute cosine similarity (cosine distance = 1 - similarity)
cos_sim <- proxy::simil(tfidf_mat, method = "cosine")
cos_sim_mat <- as.matrix(cos_sim)

# Set a similarity threshold (adjust based on your needs; 0.9 works well here)
threshold <- 0.9

# Identify which rows to keep (longest text in each duplicate group)
keep_indices <- c()
for (i in 1:nrow(data)) {
  if (!i %in% keep_indices) {
    # Find all rows highly similar to the current one
    similar_rows <- which(cos_sim_mat[i, ] >= threshold)
    # Pick the row with the longest text
    longest_row <- similar_rows[which.max(nchar(data$text[similar_rows]))]
    keep_indices <- c(keep_indices, longest_row)
  }
}

# Final deduplicated dataset
deduplicated_data <- data[keep_indices, ]

4. Optional: Simplify with your existing ID field

Notice your ID column already groups related entries (e.g., Foot123 for both football texts). If this ID is reliable, you can skip the cosine distance step entirely and just keep the longest text per ID:

deduplicated_data_by_id <- data %>%
  group_by(ID) %>%
  slice(which.max(nchar(text))) %>%
  ungroup()

This is faster if your ID field accurately maps to duplicate groups—use this first if you trust the ID labels!

Alternative approaches to consider

  • For very large datasets (millions of rows), cosine distance can be slow. Use SimHash (a locality-sensitive hash algorithm) to find near-duplicates much faster.
  • For short texts, Jaccard similarity works well, but TF-IDF + cosine distance is better for long paragraphs like yours.

内容的提问来源于stack exchange,提问作者user8959427

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:17:32