如何在Python中验证大文本的余弦、杰卡德等相似度指标准确性
Hey there! Let's break down how to tackle your text analysis project step by step—we'll cover validating your existing similarity metrics first, then implement the two new ones you need, and wrap it all up with integrating everything into a pandas DataFrame.
It makes sense to be cautious about large file results when small tests worked—let's narrow down why discrepancies might happen and verify accuracy:
Run sanity checks with controlled text snippets
Pull small, intentional chunks from your large files (e.g., a near-identical pair, a completely unrelated pair, and a partially overlapping pair) and calculate the metrics manually. Compare those manual results to what your code outputs—this will quickly reveal if your implementation has a logical flaw.Audit text preprocessing consistency
Large texts often have more noise (special characters, inconsistent casing, redundant whitespace) that can skew results. Ensure both reports go through identical preprocessing steps before similarity calculation:def preprocess_text(text): # Lowercase all text text = text.lower() # Remove punctuation (adjust based on your data) text = text.replace('.', '').replace(',', '').replace('!', '').replace('?', '') # Optional: Filter stopwords (use nltk or scikit-learn) from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS words = [word for word in text.split() if word not in ENGLISH_STOP_WORDS] return ' '.join(words)Verify Jaccard implementation details
Jaccard similarity is strictly based on word sets (ignores word frequency)—if your code accidentally uses word counts instead of unique words, results for large texts will be wrong. Here's a correct reference implementation:def jaccard_similarity(text1, text2): words1 = set(text1.split()) words2 = set(text2.split()) intersection = len(words1 & words2) union = len(words1 | words2) return intersection / union if union != 0 else 0.0Cross-validate with trusted libraries
Use scikit-learn's built-in functions to cross-check your results. For example:from sklearn.metrics.pairwise import cosine_similarity from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.metrics import jaccard_score import numpy as np # Validate cosine similarity vectorizer = TfidfVectorizer(stop_words='english') tfidf_matrix = vectorizer.fit_transform([preprocess_text(text1), preprocess_text(text2)]) library_cosine = cosine_similarity(tfidf_matrix[0:1], tfidf_matrix[1:2])[0][0] # Validate Jaccard similarity word_set = set(vectorizer.get_feature_names_out()) text1_binary = np.array([1 if word in preprocess_text(text1).split() else 0 for word in word_set]) text2_binary = np.array([1 if word in preprocess_text(text2).split() else 0 for word in word_set]) library_jaccard = jaccard_score(text1_binary, text2_binary)Compare these library results to your own code's output—any big differences mean you need to debug your implementation.
Sim_MinEdit is typically based on normalized Levenshtein distance (edit distance divided by the longest text length, subtracted from 1 to get a similarity score between 0 and 1). For efficiency with large texts, use a dedicated library:
import Levenshtein # Install first with: pip install python-Levenshtein def sim_minedit(text1, text2): edit_distance = Levenshtein.distance(text1, text2) max_length = max(len(text1), len(text2)) if max_length == 0: return 1.0 # Edge case: empty texts return 1 - (edit_distance / max_length)
If you can't use external libraries, you can implement a dynamic programming version of Levenshtein distance—but note it will be slower for large texts.
Assuming Sim_Simple refers to a word frequency-weighted overlap metric (a common "simple" similarity measure), here's a robust implementation:
from collections import Counter def sim_simple(text1, text2): words1 = text1.split() words2 = text2.split() if not words1 or not words2: return 0.0 count1 = Counter(words1) count2 = Counter(words2) # Sum of minimum frequencies for overlapping words intersection_sum = sum(min(count1[word], count2[word]) for word in count1 if word in count2) # Sum of total frequencies minus intersection to avoid double-counting union_sum = sum(count1.values()) + sum(count2.values()) - intersection_sum return intersection_sum / union_sum if union_sum != 0 else 0.0
If your definition of Sim_Simple differs (e.g., longest common substring length ratio), adjust the logic accordingly—this version is intuitive and works well for most text comparison use cases.
Finally, wrap everything together to process your folder of reports and store results in a DataFrame:
import os import pandas as pd def compute_all_similarities(text1, text2): t1 = preprocess_text(text1) t2 = preprocess_text(text2) return { 'cosine': cosine_similarity(t1, t2), # Replace with your cosine implementation 'jaccard': jaccard_similarity(t1, t2), 'minedit': sim_minedit(t1, t2), 'simple': sim_simple(t1, t2) } # Configure your folder path folder_path = '/path/to/your/reports' files = [f for f in os.listdir(folder_path) if os.path.isfile(os.path.join(folder_path, f))] # Calculate similarities for all file pairs (adjust if you need 1-to-1 comparisons instead of pairs) results = [] for i in range(len(files)): for j in range(i + 1, len(files)): file1_path = os.path.join(folder_path, files[i]) file2_path = os.path.join(folder_path, files[j]) with open(file1_path, 'r', encoding='utf-8') as f: text1 = f.read() with open(file2_path, 'r', encoding='utf-8') as f: text2 = f.read() sim_scores = compute_all_similarities(text1, text2) results.append({ 'file_1': files[i], 'file_2': files[j], **sim_scores }) # Convert to DataFrame and save similarity_df = pd.DataFrame(results) print(similarity_df) similarity_df.to_csv('report_similarities.csv', index=False)
内容的提问来源于stack exchange,提问作者ibarant

