You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在词语相似度计算中为前缀分配更高权重?(已知Jaro-Winkler)

Word-Level Similarity Methods with Prefix Priority

Great question! Since Jaro-Winkler only operates at the character level, you're right to seek word-level approaches that prioritize prefix terms (like the "X-ray" or "Rgraphy" at the start of your medical phrases). Here are three practical methods tailored to your use case, with actionable examples you can adapt to your R data frames:

1. Weighted Bag-of-Words (BoW) with Position-Based Prefix Weights

This is a straightforward, easy-to-implement approach: split each phrase into individual words, assign higher weights to words that appear earlier in the sequence, then calculate similarity using weighted vector comparisons (like cosine similarity).

For example, you could set weights like:

  • 3x weight to the 1st word (your core prefix)
  • 2x weight to the 2nd word
  • 1x weight to all subsequent words

Here’s how to build this in R using tidytext and proxy:

library(tidytext)
library(dplyr)
library(proxy)

# Helper function to tokenize phrases with position and weight
tokenize_with_weights <- function(df) {
  df %>%
    mutate(id = row_number()) %>%
    unnest_tokens(word, name) %>%
    group_by(id) %>%
    mutate(
      position = row_number(),
      weight = case_when(
        position == 1 ~ 3,
        position == 2 ~ 2,
        TRUE ~ 1
      )
    ) %>%
    ungroup()
}

# Process your A and B data frames
A_tokens <- tokenize_with_weights(A)
B_tokens <- tokenize_with_weights(B)

# Create weighted document-term matrices
build_weighted_dtm <- function(tokens) {
  tokens %>%
    mutate(weighted_count = 1 * weight) %>%
    cast_dtm(id, word, weighted_count)
}

A_dtm <- build_weighted_dtm(A_tokens)
B_dtm <- build_weighted_dtm(B_tokens)

# Calculate cosine similarity between all phrase pairs
similarity_matrix <- proxy::simil(A_dtm, B_dtm, method = "cosine")
print(similarity_matrix)

This will output similarity scores where matches in the prefix word have a far larger impact on the final score than matches later in the phrase.

2. Weighted Word Embeddings (Semantic Similarity)

If you want to capture semantic similarity (not just exact word matches), you can use pre-trained word embeddings (like GloVe or Word2Vec) and amplify the weight of prefix words in the embedding calculation.

The workflow is:

  1. Fetch embedding vectors for each word in the phrase
  2. Multiply vectors of the first 1-2 words by a weight (e.g., 2x)
  3. Average all weighted vectors to get a phrase-level embedding
  4. Compute cosine similarity between phrase embeddings

Here’s a simplified example using text2vec:

library(text2vec)

# Load pre-trained GloVe embeddings (you can download free public versions)
glove <- readRDS("glove.6B.100d.rds")

# Function to generate weighted phrase embeddings
get_weighted_embedding <- function(phrase) {
  words <- str_split(phrase, " ")[[1]]
  # Filter words present in the embedding set
  valid_words <- words[words %in% rownames(glove)]
  if (length(valid_words) == 0) return(rep(0, 100))
  
  # Assign prefix weights
  weights <- ifelse(seq_along(words) == 1, 2, 
                    ifelse(seq_along(words) == 2, 1.5, 1))
  weights <- weights[words %in% rownames(glove)]
  
  # Calculate weighted average embedding
  colSums(glove[valid_words, , drop = FALSE] * weights) / sum(weights)
}

# Generate embeddings for all phrases in A and B
A_embeds <- sapply(A$name, get_weighted_embedding)
B_embeds <- sapply(B$name, get_weighted_embedding)

# Compute similarity matrix
similarity_matrix <- proxy::simil(t(A_embeds), t(B_embeds), method = "cosine")
print(similarity_matrix)

3. Custom Sequence Matching with Prefix Bonuses

For full control over scoring, build a custom function that gives extra points for matching words in prefix positions. For example:

  • +5 points for exact matches in the first word
  • +3 points for exact matches in the second word
  • +1 point for exact matches in later words
  • Normalize by total possible points to get a 0-1 similarity score

Here’s a quick R implementation:

custom_prefix_similarity <- function(phrase1, phrase2) {
  words1 <- str_split(tolower(phrase1), " ")[[1]]
  words2 <- str_split(tolower(phrase2), " ")[[1]]
  
  max_len <- max(length(words1), length(words2))
  score <- 0
  total_possible <- 0
  
  for (i in 1:max_len) {
    # Assign position-based weights
    weight <- ifelse(i == 1, 5, ifelse(i == 2, 3, 1))
    total_possible <- total_possible + weight
    
    # Check for matching words
    if (i <= length(words1) && i <= length(words2)) {
      if (words1[i] == words2[i]) score <- score + weight
    }
  }
  
  score / total_possible
}

# Apply to all phrase pairs
similarity_matrix <- outer(A$name, B$name, Vectorize(custom_prefix_similarity))
rownames(similarity_matrix) <- A$name
colnames(similarity_matrix) <- B$name
print(similarity_matrix)

All three methods prioritize prefix matches in word-level similarity calculations—pick the one that best balances simplicity and your need for semantic depth!

内容的提问来源于stack exchange,提问作者john

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:27:46