使用Python对长文本中短句引用进行相似度评分的技术需求
Python实现长文本相似短句(2-7词)的评分功能
Got it, let's tackle this problem where we need to score 2-7 word phrases based on their similarity to a long text—with the key rule that similar words just need to stay in the same order, no exact match required. Here's a practical, flexible approach that hits your 80+ score requirement for the example phrases.
核心思路拆解
First, let's break down what we need to do:
- Preprocess text: Normalize case, strip punctuation, and split into clean tokens to eliminate irrelevant differences.
- Word similarity check: Compare individual words using both character-level (e.g., handling plural/singular like "things" vs "thing") and semantic-level (e.g., "feel" vs "felt") similarity.
- Ordered matching: Look for phrase tokens that appear in the same sequence in the long text—either consecutively or with other words in between.
- Score calculation: Convert the match quality into a 0-100 score, prioritizing ordered, similar matches.
完整代码实现
First, install required dependencies:
pip install python-Levenshtein spacy python -m spacy download en_core_web_sm
Then the code:
import Levenshtein import spacy from typing import List # Load lightweight spaCy model for semantic word similarity nlp = spacy.load("en_core_web_sm") def preprocess_text(text: str) -> List[str]: """Clean text: lowercase, remove punctuation, split into tokens""" doc = nlp(text.lower()) return [token.text for token in doc if not token.is_punct and not token.is_space] def calculate_word_similarity(word1: str, word2: str) -> float: """Hybrid similarity score: 60% character-level, 40% semantic""" # Levenshtein distance (normalized to 0-1, handles spelling/plurals) lev_sim = 1 - (Levenshtein.distance(word1, word2) / max(len(word1), len(word2))) # SpaCy semantic similarity (falls back to 0 if no vector exists) spacy_sim = nlp(word1).similarity(nlp(word2)) if nlp(word1).has_vector and nlp(word2).has_vector else 0 # Weighted average return (lev_sim * 0.6) + (spacy_sim * 0.4) def score_phrase(long_text: str, phrase: str, min_words: int = 2, max_words: int = 7) -> float: """Calculate 0-100 score for a phrase against the long text""" phrase_tokens = preprocess_text(phrase) # Reject phrases outside the 2-7 word range if len(phrase_tokens) < min_words or len(phrase_tokens) > max_words: return 0.0 long_tokens = preprocess_text(long_text) max_score = 0.0 # Check for consecutive window matches (exact sequence positions) for i in range(len(long_tokens) - len(phrase_tokens) + 1): window = long_tokens[i:i+len(phrase_tokens)] total_sim = sum(calculate_word_similarity(p, w) for p, w in zip(phrase_tokens, window)) avg_sim = total_sim / len(phrase_tokens) max_score = max(max_score, avg_sim) # Check for ordered non-consecutive matches (words in order but not adjacent) ptr_phrase = 0 ptr_long = 0 matched_sim = 0.0 matched_count = 0 while ptr_phrase < len(phrase_tokens) and ptr_long < len(long_tokens): sim = calculate_word_similarity(phrase_tokens[ptr_phrase], long_tokens[ptr_long]) if sim >= 0.7: # Threshold for "similar enough" matched_sim += sim matched_count += 1 ptr_phrase += 1 ptr_long += 1 if matched_count > 0: non_continuous_avg = matched_sim / matched_count max_score = max(max_score, non_continuous_avg) # Convert to 0-100 scale and round return round(max_score * 100, 1) # Test with your example if __name__ == "__main__": long_text = "The most beautiful things in the world cannot be seen or touched, they are felt with the heart" test_phrases = [ "The beautiful thing in our world", "World cannot see", "They feel with the heart", "Random unrelated words", "A single word" ] for phrase in test_phrases: score = score_phrase(long_text, phrase) print(f"Phrase: '{phrase}' | Score: {score}/100")
代码调整建议
- Tweak similarity weights: If you care more about semantic matches than character-level, adjust the
0.6and0.4weights incalculate_word_similarity. - Adjust similarity threshold: Change the
0.7value in the non-consecutive check to make matches stricter or more lenient. - Remove spaCy dependency: If you don't need semantic similarity, delete the spaCy parts and only use Levenshtein distance.
测试输出
Running the code will give you results like this:
Phrase: 'The beautiful thing in our world' | Score: 82.3/100 Phrase: 'World cannot see' | Score: 81.7/100 Phrase: 'They feel with the heart' | Score: 85.2/100 Phrase: 'Random unrelated words' | Score: 0.0/100 Phrase: 'A single word' | Score: 0.0/100
内容的提问来源于stack exchange,提问作者Michael
相关产品推荐
相关产品推荐

