You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中基于模型句子查找文本内相似句子的方案咨询

Got it, let's tackle this problem step by step. You want to find sentences in your data dataset that are similar to the model sentences in model_sentences—here are two practical approaches in R, starting with a simple, fast method and moving to a more powerful semantic similarity solution.

First, let's recap your datasets for clarity:

model_sentences <- data.frame(
  "model_id" = c("model_id_1", "model_id_2"),
  "model_text" = c("Company x had 3000 employees in 2016.", "Google makes 300 dollar in revenue in 2018.")
)

data <- data.frame(
  "id" = c("id1", "id2"),
  "text" = c("Company y is expected to employ 2000 employees in 2020. This is an increase of 10%. Some stupid sentences.", "Amazon´s revenue is 400...")
)

Approach 1: TF-IDF + Cosine Similarity (Quick & Simple)

This method works well for matching sentences with overlapping vocabulary, and it's lightweight enough for small to medium datasets.

Step 1: Install & Load Required Packages

library(dplyr)
library(stringr)
library(tidytext)
library(proxy)

Step 2: Split Long Texts into Individual Sentences

Your data$text has multi-sentence entries, so first we'll split them into separate rows:

data_sentences <- data %>%
  # Split on punctuation followed by whitespace
  mutate(sentence = str_split(text, "(?<=[.!?])\\s+")) %>%
  unnest(sentence) %>%
  filter(str_trim(sentence) != "") # Remove empty sentences

Step 3: Build a Combined Corpus & Calculate TF-IDF

We'll combine model sentences and candidate sentences to create a unified TF-IDF matrix:

# Merge model and candidate sentences
corpus <- bind_rows(
  model_sentences %>% mutate(type = "model") %>% rename(sentence = model_text),
  data_sentences %>% mutate(type = "candidate") %>% select(sentence, id, type)
)

# Generate TF-IDF matrix
tf_idf <- corpus %>%
  unnest_tokens(word, sentence) %>%
  count(type, id = ifelse(type == "model", model_id, id), word) %>%
  bind_tf_idf(word, id, n) %>%
  pivot_wider(names_from = word, values_from = tf_idf, values_fill = 0)

# Separate model and candidate matrices
model_tfidf <- tf_idf %>% filter(type == "model") %>% select(-type, -id)
candidate_tfidf <- tf_idf %>% filter(type == "candidate") %>% select(-type, -id)
candidate_ids <- tf_idf %>% filter(type == "candidate") %>% pull(id)

Step 4: Calculate Cosine Similarity

Cosine similarity measures how "close" two vectors (sentences) are in the TF-IDF space:

# Compute similarity matrix
similarity_matrix <- proxy::simil(candidate_tfidf, model_tfidf, method = "cosine")

# Convert to readable dataframe
similarity_results <- similarity_matrix %>%
  as.data.frame() %>%
  mutate(candidate_id = candidate_ids) %>%
  rename_with(~ paste0("similarity_to_", .x), starts_with("model_id")) %>%
  relocate(candidate_id)

Step 5: Filter Top Matches

Set a similarity threshold (e.g., > 0.5) to get relevant matches:

top_matches <- similarity_results %>%
  filter(if_any(starts_with("similarity_to"), ~ .x > 0.5))

Pros: Fast, no external model dependencies.
Cons: Struggles with semantic similarity where vocabulary differs (e.g., "employ" vs "had...employees").


Approach 2: BERT Sentence Embeddings (Semantic Similarity)

For better results with meaning-based matches (even if words differ), use a pre-trained language model like BERT. This captures the context of sentences, so it'll recognize that "Company y will employ 2000 employees" is similar to "Company x had 3000 employees".

Step 1: Install & Load BERT Package

# Install if needed
# remotes::install_github("jcrodriguez1989/bert4r")

library(bert4r)
library(purrr)

Step 2: Load Pre-trained BERT Model

We'll use the lightweight bert-base-uncased model:

model <- load_bert("bert-base-uncased")

Step 3: Generate Sentence Embeddings

Embeddings are numerical representations of sentence meaning:

# Get embeddings for model sentences
model_embeddings <- model_sentences %>%
  mutate(embedding = map(model_text, ~ get_embeddings(model, .x)$cls_embedding))

# Get embeddings for candidate sentences
candidate_embeddings <- data_sentences %>%
  mutate(embedding = map(sentence, ~ get_embeddings(model, .x)$cls_embedding))

Step 4: Calculate Semantic Similarity

# Helper function to compute similarity between a candidate and all model sentences
calculate_similarity <- function(cand_emb, model_embs) {
  map_dbl(model_embs, ~ proxy::simil(matrix(cand_emb, nrow=1), matrix(.x, nrow=1), method="cosine"))
}

# Compute and format results
similarity_results_bert <- candidate_embeddings %>%
  mutate(similarities = map(embedding, calculate_similarity, model_embs = model_embeddings$embedding)) %>%
  unnest_wider(similarities, names_sep = "_") %>%
  rename_with(~ paste0("similarity_to_model_id_", str_remove(.x, "similarities_")), starts_with("similarities"))

Pros: Excellent at capturing semantic similarity, works well with domain-specific text.
Cons: Requires more computational resources, first model load takes time.


Key Notes & Adjustments
  • Preprocessing: For both methods, you can add steps like lowercasing text, removing stopwords, or cleaning punctuation to improve results.
  • Threshold Tuning: Adjust the similarity threshold based on your use case—higher thresholds mean stricter matches.
  • Model Choice: For BERT, you can use domain-specific models (e.g., bert-base-financial-uncased) if your text is about finance, which will improve accuracy.

内容的提问来源于stack exchange,提问作者WinterMensch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:30:22