如何在JavaScript中提取两段文本间符合规则的相似句子序列
Got it, let's break down how to build this algorithm step by step—this is a classic text similarity problem with specific constraints, so we'll tailor the approach to your exact rules.
1. Preprocess: Split Texts into Individual Sentences
First, you need to reliably split both input texts into complete sentences. Forget naive splitting by periods (they show up in abbreviations like "Mr." or "etc."). Use battle-tested NLP tools instead:
- For Python: Use
nltk.tokenize.sent_tokenizeor spaCy's built-in sentencizer. - For other languages: Look for equivalent libraries (e.g., CoreNLP for Java, spaCy for JavaScript).
2. Identify Continuous Matching Sentence Groups
Your requirement focuses on continuous sequences of full sentences (not scattered one-off matches). A dynamic programming (DP) approach works perfectly here to find the longest (and all valid) continuous matching sequences:
- Use a 2D DP array where
dp[i][j]represents the length of the longest continuous matching sequence ending at the i-th sentence of text1 and j-th sentence of text2. - For exact matches: Check if the current sentences are identical (after stripping whitespace and normalizing case if needed).
- For semantic matches (if you need to account for paraphrasing): Use sentence embeddings (like Sentence-BERT) to calculate similarity scores, then set a threshold (e.g., 0.9) to define a "match".
3. Filter Out Common Daily Phrases
Create a curated list of overused, generic phrases you want to ignore. Then, validate each matching group:
- Remove any sentence that falls into this common phrases list.
- Only keep the group if it still contains at least one valid sentence after filtering.
Full Python Implementation Example
import nltk from sentence_transformers import SentenceTransformer, util # Download required NLTK data for sentence tokenization nltk.download('punkt') from nltk.tokenize import sent_tokenize # Initialize a lightweight sentence embedding model for semantic matching model = SentenceTransformer('all-MiniLM-L6-v2') # Your input texts text1 = "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est here." text2 = "Hi, I'm looking for a course to learn programming. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat idk what I'm supposed to write down here." # Step 1: Split texts into individual sentences sentences1 = sent_tokenize(text1) sentences2 = sent_tokenize(text2) # Step 2: Find continuous matching groups (semantic match example) similarity_threshold = 0.9 dp = [[0]*(len(sentences2)+1) for _ in range(len(sentences1)+1)] max_match_length = 0 match_positions = [] # Compute sentence embeddings for similarity checks embeddings1 = model.encode(sentences1, convert_to_tensor=True) embeddings2 = model.encode(sentences2, convert_to_tensor=True) cosine_scores = util.cos_sim(embeddings1, embeddings2) # Fill the DP table to track continuous matches for i in range(1, len(sentences1)+1): for j in range(1, len(sentences2)+1): if cosine_scores[i-1][j-1] >= similarity_threshold: dp[i][j] = dp[i-1][j-1] + 1 if dp[i][j] > max_match_length: max_match_length = dp[i][j] match_positions = [(i - max_match_length, i)] elif dp[i][j] == max_match_length and max_match_length >= 1: match_positions.append((i - max_match_length, i)) # Step 3: Filter out common daily phrases common_phrases = { "Hope to hear from you soon.", "God bless you.", "Thank you very much.", "Best regards.", # Add more generic phrases as needed } # Extract and validate matches valid_matches = [] for start_idx, end_idx in match_positions: sentence_group = sentences1[start_idx:end_idx] # Filter out any common phrases from the group filtered_group = [sent.strip() for sent in sentence_group if sent.strip() not in common_phrases] if filtered_group: valid_matches.append(' '.join(filtered_group)) # Output the final results print("Valid matching sequences:") for match in valid_matches: print(f"> {match}")
Notes for Customization
- Exact vs. Semantic Matching: If you only need strict exact matches, replace the cosine similarity check with a direct string comparison (e.g.,
sentences1[i-1].strip().lower() == sentences2[j-1].strip().lower()). - Handling Shorter Matches: The code above collects only the longest matching groups. If you want all valid sequences (even shorter ones), adjust the DP logic to track all sequences of length ≥1.
- Expanding Common Phrases: You can grow the
common_phrasesset with more entries, or use a pre-built generic phrase corpus from NLP datasets if you need a larger list.
内容的提问来源于stack exchange,提问作者Efe FRK

