You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在JavaScript中提取两段文本间符合规则的相似句子序列

Approach to Extract Matching Full Sentence Groups Between Two Texts

Got it, let's break down how to build this algorithm step by step—this is a classic text similarity problem with specific constraints, so we'll tailor the approach to your exact rules.

1. Preprocess: Split Texts into Individual Sentences

First, you need to reliably split both input texts into complete sentences. Forget naive splitting by periods (they show up in abbreviations like "Mr." or "etc."). Use battle-tested NLP tools instead:

  • For Python: Use nltk.tokenize.sent_tokenize or spaCy's built-in sentencizer.
  • For other languages: Look for equivalent libraries (e.g., CoreNLP for Java, spaCy for JavaScript).

2. Identify Continuous Matching Sentence Groups

Your requirement focuses on continuous sequences of full sentences (not scattered one-off matches). A dynamic programming (DP) approach works perfectly here to find the longest (and all valid) continuous matching sequences:

  • Use a 2D DP array where dp[i][j] represents the length of the longest continuous matching sequence ending at the i-th sentence of text1 and j-th sentence of text2.
  • For exact matches: Check if the current sentences are identical (after stripping whitespace and normalizing case if needed).
  • For semantic matches (if you need to account for paraphrasing): Use sentence embeddings (like Sentence-BERT) to calculate similarity scores, then set a threshold (e.g., 0.9) to define a "match".

3. Filter Out Common Daily Phrases

Create a curated list of overused, generic phrases you want to ignore. Then, validate each matching group:

  • Remove any sentence that falls into this common phrases list.
  • Only keep the group if it still contains at least one valid sentence after filtering.

Full Python Implementation Example

import nltk
from sentence_transformers import SentenceTransformer, util

# Download required NLTK data for sentence tokenization
nltk.download('punkt')
from nltk.tokenize import sent_tokenize

# Initialize a lightweight sentence embedding model for semantic matching
model = SentenceTransformer('all-MiniLM-L6-v2')

# Your input texts
text1 = "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est here."
text2 = "Hi, I'm looking for a course to learn programming. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat idk what I'm supposed to write down here."

# Step 1: Split texts into individual sentences
sentences1 = sent_tokenize(text1)
sentences2 = sent_tokenize(text2)

# Step 2: Find continuous matching groups (semantic match example)
similarity_threshold = 0.9
dp = [[0]*(len(sentences2)+1) for _ in range(len(sentences1)+1)]
max_match_length = 0
match_positions = []

# Compute sentence embeddings for similarity checks
embeddings1 = model.encode(sentences1, convert_to_tensor=True)
embeddings2 = model.encode(sentences2, convert_to_tensor=True)
cosine_scores = util.cos_sim(embeddings1, embeddings2)

# Fill the DP table to track continuous matches
for i in range(1, len(sentences1)+1):
    for j in range(1, len(sentences2)+1):
        if cosine_scores[i-1][j-1] >= similarity_threshold:
            dp[i][j] = dp[i-1][j-1] + 1
            if dp[i][j] > max_match_length:
                max_match_length = dp[i][j]
                match_positions = [(i - max_match_length, i)]
            elif dp[i][j] == max_match_length and max_match_length >= 1:
                match_positions.append((i - max_match_length, i))

# Step 3: Filter out common daily phrases
common_phrases = {
    "Hope to hear from you soon.",
    "God bless you.",
    "Thank you very much.",
    "Best regards.",
    # Add more generic phrases as needed
}

# Extract and validate matches
valid_matches = []
for start_idx, end_idx in match_positions:
    sentence_group = sentences1[start_idx:end_idx]
    # Filter out any common phrases from the group
    filtered_group = [sent.strip() for sent in sentence_group if sent.strip() not in common_phrases]
    if filtered_group:
        valid_matches.append(' '.join(filtered_group))

# Output the final results
print("Valid matching sequences:")
for match in valid_matches:
    print(f"> {match}")

Notes for Customization

  • Exact vs. Semantic Matching: If you only need strict exact matches, replace the cosine similarity check with a direct string comparison (e.g., sentences1[i-1].strip().lower() == sentences2[j-1].strip().lower()).
  • Handling Shorter Matches: The code above collects only the longest matching groups. If you want all valid sequences (even shorter ones), adjust the DP logic to track all sequences of length ≥1.
  • Expanding Common Phrases: You can grow the common_phrases set with more entries, or use a pre-built generic phrase corpus from NLP datasets if you need a larger list.

内容的提问来源于stack exchange,提问作者Efe FRK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 04:24:06