R语言中基于模型句子查找文本内相似句子的方案咨询
Got it, let's tackle this problem step by step. You want to find sentences in your data dataset that are similar to the model sentences in model_sentences—here are two practical approaches in R, starting with a simple, fast method and moving to a more powerful semantic similarity solution.
First, let's recap your datasets for clarity:
model_sentences <- data.frame( "model_id" = c("model_id_1", "model_id_2"), "model_text" = c("Company x had 3000 employees in 2016.", "Google makes 300 dollar in revenue in 2018.") ) data <- data.frame( "id" = c("id1", "id2"), "text" = c("Company y is expected to employ 2000 employees in 2020. This is an increase of 10%. Some stupid sentences.", "Amazon´s revenue is 400...") )
This method works well for matching sentences with overlapping vocabulary, and it's lightweight enough for small to medium datasets.
Step 1: Install & Load Required Packages
library(dplyr) library(stringr) library(tidytext) library(proxy)
Step 2: Split Long Texts into Individual Sentences
Your data$text has multi-sentence entries, so first we'll split them into separate rows:
data_sentences <- data %>% # Split on punctuation followed by whitespace mutate(sentence = str_split(text, "(?<=[.!?])\\s+")) %>% unnest(sentence) %>% filter(str_trim(sentence) != "") # Remove empty sentences
Step 3: Build a Combined Corpus & Calculate TF-IDF
We'll combine model sentences and candidate sentences to create a unified TF-IDF matrix:
# Merge model and candidate sentences corpus <- bind_rows( model_sentences %>% mutate(type = "model") %>% rename(sentence = model_text), data_sentences %>% mutate(type = "candidate") %>% select(sentence, id, type) ) # Generate TF-IDF matrix tf_idf <- corpus %>% unnest_tokens(word, sentence) %>% count(type, id = ifelse(type == "model", model_id, id), word) %>% bind_tf_idf(word, id, n) %>% pivot_wider(names_from = word, values_from = tf_idf, values_fill = 0) # Separate model and candidate matrices model_tfidf <- tf_idf %>% filter(type == "model") %>% select(-type, -id) candidate_tfidf <- tf_idf %>% filter(type == "candidate") %>% select(-type, -id) candidate_ids <- tf_idf %>% filter(type == "candidate") %>% pull(id)
Step 4: Calculate Cosine Similarity
Cosine similarity measures how "close" two vectors (sentences) are in the TF-IDF space:
# Compute similarity matrix similarity_matrix <- proxy::simil(candidate_tfidf, model_tfidf, method = "cosine") # Convert to readable dataframe similarity_results <- similarity_matrix %>% as.data.frame() %>% mutate(candidate_id = candidate_ids) %>% rename_with(~ paste0("similarity_to_", .x), starts_with("model_id")) %>% relocate(candidate_id)
Step 5: Filter Top Matches
Set a similarity threshold (e.g., > 0.5) to get relevant matches:
top_matches <- similarity_results %>% filter(if_any(starts_with("similarity_to"), ~ .x > 0.5))
Pros: Fast, no external model dependencies.
Cons: Struggles with semantic similarity where vocabulary differs (e.g., "employ" vs "had...employees").
For better results with meaning-based matches (even if words differ), use a pre-trained language model like BERT. This captures the context of sentences, so it'll recognize that "Company y will employ 2000 employees" is similar to "Company x had 3000 employees".
Step 1: Install & Load BERT Package
# Install if needed # remotes::install_github("jcrodriguez1989/bert4r") library(bert4r) library(purrr)
Step 2: Load Pre-trained BERT Model
We'll use the lightweight bert-base-uncased model:
model <- load_bert("bert-base-uncased")
Step 3: Generate Sentence Embeddings
Embeddings are numerical representations of sentence meaning:
# Get embeddings for model sentences model_embeddings <- model_sentences %>% mutate(embedding = map(model_text, ~ get_embeddings(model, .x)$cls_embedding)) # Get embeddings for candidate sentences candidate_embeddings <- data_sentences %>% mutate(embedding = map(sentence, ~ get_embeddings(model, .x)$cls_embedding))
Step 4: Calculate Semantic Similarity
# Helper function to compute similarity between a candidate and all model sentences calculate_similarity <- function(cand_emb, model_embs) { map_dbl(model_embs, ~ proxy::simil(matrix(cand_emb, nrow=1), matrix(.x, nrow=1), method="cosine")) } # Compute and format results similarity_results_bert <- candidate_embeddings %>% mutate(similarities = map(embedding, calculate_similarity, model_embs = model_embeddings$embedding)) %>% unnest_wider(similarities, names_sep = "_") %>% rename_with(~ paste0("similarity_to_model_id_", str_remove(.x, "similarities_")), starts_with("similarities"))
Pros: Excellent at capturing semantic similarity, works well with domain-specific text.
Cons: Requires more computational resources, first model load takes time.
- Preprocessing: For both methods, you can add steps like lowercasing text, removing stopwords, or cleaning punctuation to improve results.
- Threshold Tuning: Adjust the similarity threshold based on your use case—higher thresholds mean stricter matches.
- Model Choice: For BERT, you can use domain-specific models (e.g.,
bert-base-financial-uncased) if your text is about finance, which will improve accuracy.
内容的提问来源于stack exchange,提问作者WinterMensch

