You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中验证大文本的余弦、杰卡德等相似度指标准确性

Hey there! Let's break down how to tackle your text analysis project step by step—we'll cover validating your existing similarity metrics first, then implement the two new ones you need, and wrap it all up with integrating everything into a pandas DataFrame.

1. Validating Cosine & Jaccard Similarity for Large Texts

It makes sense to be cautious about large file results when small tests worked—let's narrow down why discrepancies might happen and verify accuracy:

  • Run sanity checks with controlled text snippets
    Pull small, intentional chunks from your large files (e.g., a near-identical pair, a completely unrelated pair, and a partially overlapping pair) and calculate the metrics manually. Compare those manual results to what your code outputs—this will quickly reveal if your implementation has a logical flaw.

  • Audit text preprocessing consistency
    Large texts often have more noise (special characters, inconsistent casing, redundant whitespace) that can skew results. Ensure both reports go through identical preprocessing steps before similarity calculation:

    def preprocess_text(text):
        # Lowercase all text
        text = text.lower()
        # Remove punctuation (adjust based on your data)
        text = text.replace('.', '').replace(',', '').replace('!', '').replace('?', '')
        # Optional: Filter stopwords (use nltk or scikit-learn)
        from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
        words = [word for word in text.split() if word not in ENGLISH_STOP_WORDS]
        return ' '.join(words)
    
  • Verify Jaccard implementation details
    Jaccard similarity is strictly based on word sets (ignores word frequency)—if your code accidentally uses word counts instead of unique words, results for large texts will be wrong. Here's a correct reference implementation:

    def jaccard_similarity(text1, text2):
        words1 = set(text1.split())
        words2 = set(text2.split())
        intersection = len(words1 & words2)
        union = len(words1 | words2)
        return intersection / union if union != 0 else 0.0
    
  • Cross-validate with trusted libraries
    Use scikit-learn's built-in functions to cross-check your results. For example:

    from sklearn.metrics.pairwise import cosine_similarity
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.metrics import jaccard_score
    import numpy as np
    
    # Validate cosine similarity
    vectorizer = TfidfVectorizer(stop_words='english')
    tfidf_matrix = vectorizer.fit_transform([preprocess_text(text1), preprocess_text(text2)])
    library_cosine = cosine_similarity(tfidf_matrix[0:1], tfidf_matrix[1:2])[0][0]
    
    # Validate Jaccard similarity
    word_set = set(vectorizer.get_feature_names_out())
    text1_binary = np.array([1 if word in preprocess_text(text1).split() else 0 for word in word_set])
    text2_binary = np.array([1 if word in preprocess_text(text2).split() else 0 for word in word_set])
    library_jaccard = jaccard_score(text1_binary, text2_binary)
    

    Compare these library results to your own code's output—any big differences mean you need to debug your implementation.

2. Implementing Sim_MinEdit (Edit Distance Similarity)

Sim_MinEdit is typically based on normalized Levenshtein distance (edit distance divided by the longest text length, subtracted from 1 to get a similarity score between 0 and 1). For efficiency with large texts, use a dedicated library:

import Levenshtein  # Install first with: pip install python-Levenshtein

def sim_minedit(text1, text2):
    edit_distance = Levenshtein.distance(text1, text2)
    max_length = max(len(text1), len(text2))
    if max_length == 0:
        return 1.0  # Edge case: empty texts
    return 1 - (edit_distance / max_length)

If you can't use external libraries, you can implement a dynamic programming version of Levenshtein distance—but note it will be slower for large texts.

3. Implementing Sim_Simple (Simple Similarity)

Assuming Sim_Simple refers to a word frequency-weighted overlap metric (a common "simple" similarity measure), here's a robust implementation:

from collections import Counter

def sim_simple(text1, text2):
    words1 = text1.split()
    words2 = text2.split()
    if not words1 or not words2:
        return 0.0
    
    count1 = Counter(words1)
    count2 = Counter(words2)
    
    # Sum of minimum frequencies for overlapping words
    intersection_sum = sum(min(count1[word], count2[word]) for word in count1 if word in count2)
    # Sum of total frequencies minus intersection to avoid double-counting
    union_sum = sum(count1.values()) + sum(count2.values()) - intersection_sum
    
    return intersection_sum / union_sum if union_sum != 0 else 0.0

If your definition of Sim_Simple differs (e.g., longest common substring length ratio), adjust the logic accordingly—this version is intuitive and works well for most text comparison use cases.

4. Integrate All Metrics into a pandas DataFrame

Finally, wrap everything together to process your folder of reports and store results in a DataFrame:

import os
import pandas as pd

def compute_all_similarities(text1, text2):
    t1 = preprocess_text(text1)
    t2 = preprocess_text(text2)
    return {
        'cosine': cosine_similarity(t1, t2),  # Replace with your cosine implementation
        'jaccard': jaccard_similarity(t1, t2),
        'minedit': sim_minedit(t1, t2),
        'simple': sim_simple(t1, t2)
    }

# Configure your folder path
folder_path = '/path/to/your/reports'
files = [f for f in os.listdir(folder_path) if os.path.isfile(os.path.join(folder_path, f))]

# Calculate similarities for all file pairs (adjust if you need 1-to-1 comparisons instead of pairs)
results = []
for i in range(len(files)):
    for j in range(i + 1, len(files)):
        file1_path = os.path.join(folder_path, files[i])
        file2_path = os.path.join(folder_path, files[j])
        
        with open(file1_path, 'r', encoding='utf-8') as f:
            text1 = f.read()
        with open(file2_path, 'r', encoding='utf-8') as f:
            text2 = f.read()
        
        sim_scores = compute_all_similarities(text1, text2)
        results.append({
            'file_1': files[i],
            'file_2': files[j],
            **sim_scores
        })

# Convert to DataFrame and save
similarity_df = pd.DataFrame(results)
print(similarity_df)
similarity_df.to_csv('report_similarities.csv', index=False)

内容的提问来源于stack exchange,提问作者ibarant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:44:36