You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

文本相似度计算方法及CSV数据处理可运行代码需求问询

Solution: CSV Text Similarity Calculation with Cosine, USE, and Levenshtein

I’ve put together a complete, runnable script that reads your CSV file and computes similarity using the three methods you mentioned. Here’s everything you need to get started:

Prerequisites

First, install the required packages using pip:

pip install pandas scikit-learn tensorflow-hub python-Levenshtein tensorflow

(Note: tensorflow is required for the Universal Sentence Encoder; skip it if you already have it installed.)

Full Runnable Code

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
import tensorflow_hub as hub
from Levenshtein import distance as levenshtein_distance
import numpy as np

# Step 1: Load your CSV file
# Replace 'your_file.csv' with the actual path to your CSV
df = pd.read_csv('your_file.csv')

# Step 2: Cosine Similarity (using TF-IDF)
def compute_cosine_similarity(texts):
    # Convert text to TF-IDF vectors (weights important words more)
    vectorizer = TfidfVectorizer(stop_words='english')
    tfidf_matrix = vectorizer.fit_transform(texts)
    # Calculate pairwise cosine similarity
    return cosine_similarity(tfidf_matrix)

cos_sim_matrix = compute_cosine_similarity(df['messages'])

# Step 3: Universal Sentence Encoder (USE) Similarity
# Load Google's pre-trained USE model
use_model = hub.load("https://tfhub.dev/google/universal-sentence-encoder/4")

def compute_use_similarity(texts):
    # Generate semantic embeddings (captures meaning, not just word overlap)
    embeddings = use_model(texts)
    # Compute similarity via inner product of embeddings
    return np.inner(embeddings, embeddings)

use_sim_matrix = compute_use_similarity(df['messages'].tolist())

# Step 4: Levenshtein Similarity (normalized to 0-1 range)
def compute_levenshtein_similarity(texts):
    num_texts = len(texts)
    sim_matrix = np.zeros((num_texts, num_texts))
    
    for i in range(num_texts):
        for j in range(num_texts):
            text_a = texts[i]
            text_b = texts[j]
            max_length = max(len(text_a), len(text_b))
            
            if max_length == 0:
                sim_matrix[i][j] = 1.0  # Handle empty texts
            else:
                # Convert edit distance to a similarity score (1 = identical, 0 = no overlap)
                edit_distance = levenshtein_distance(text_a, text_b)
                sim_matrix[i][j] = 1 - (edit_distance / max_length)
    
    return sim_matrix

lev_sim_matrix = compute_levenshtein_similarity(df['messages'].tolist())

# Step 5: Print formatted results
print("=== Cosine Similarity Matrix ===")
print(pd.DataFrame(cos_sim_matrix, index=df['idx'], columns=df['idx']))

print("\n=== Universal Sentence Encoder Similarity Matrix ===")
print(pd.DataFrame(use_sim_matrix, index=df['idx'], columns=df['idx']))

print("\n=== Levenshtein Similarity Matrix ===")
print(pd.DataFrame(lev_sim_matrix, index=df['idx'], columns=df['idx']))

# Example: Compare message with idx=112 to all others
print("\n=== Comparing message idx=112 with others ===")
target_idx = df[df['idx'] == 112].index[0]
print(f"Cosine Similarity: {cos_sim_matrix[target_idx]}")
print(f"USE Similarity: {use_sim_matrix[target_idx]}")
print(f"Levenshtein Similarity: {lev_sim_matrix[target_idx]}")

Quick Explanation

Let me break down the key parts so you know what’s happening:

  1. CSV Loading: Uses pandas to read your file and keep the idx and messages columns intact.
  2. Cosine + TF-IDF: Focuses on word overlap, weighting important terms more heavily. Great for short, straightforward texts.
  3. Universal Sentence Encoder: Uses a pre-trained model to understand semantic meaning (e.g., "I have a blue car" and "My car is blue" will score highly similar).
  4. Levenshtein Similarity: Counts the number of edits needed to turn one text into another, then normalizes it to a 0-1 score. Perfect for catching typos or near-exact matches.

内容的提问来源于stack exchange,提问作者Bilgin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:15:28