You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在两个Pandas DataFrame中匹配相似文本并扩展分类标签

Matching Similar Feedback & Auto-Labeling Your Dataset

Got it, let's tackle this problem of matching your data_feed feedback entries to the reference labels in your reference dataset. Below are two practical, actionable approaches using Python—one lightweight for quick implementation, and another more accurate for nuanced semantic matching.

Approach 1: TF-IDF + Cosine Similarity (Lightweight & Fast)

This method is great for getting up and running quickly, and works well for basic short-text matching. Here's how to implement it:

Step-by-Step Implementation

  1. Load your datasets into pandas DataFrames
  2. Preprocess text to normalize it (lowercase, remove punctuation, etc.)
  3. Convert text to numerical vectors using TF-IDF
  4. Calculate cosine similarity between each data_feed entry and all reference feedbacks
  5. Attach the labels from the most similar reference entry to your data_feed
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# Load your actual data here (replace with pd.read_csv or your data source)
data_feed = pd.DataFrame({
    'feedback': [
        'Fast Delivery. Always before time.Thanks',
        'I have order brown shoe .And I got olive green shoe',
        'Delivery guy is a decent nd friendly guy',
        'Its really good .. my daughter loves it',
        'One t shirt was fully crushed rest everything is good',
        'Superfast delivery! I\'m impressed.'
    ]
})

reference = pd.DataFrame({
    'refer_feedback': [
        'The delivery was on time.',
        'he was polite enough',
        'worst products'
    ],
    'sub-category': ['delivery speed', 'delivery man behaviour', 'product quality'],
    'category': ['delivery', 'delivery', 'general'],
    'sentiment': ['positive', 'positive', 'negative']
})

# Basic text preprocessing to normalize inputs
def preprocess_text(text):
    text = text.lower().replace('.', '').replace('!', '').replace(',', '')
    # Optional: Add stopword removal (use nltk.corpus.stopwords if needed)
    return text

# Apply preprocessing to both datasets
data_feed['processed_feedback'] = data_feed['feedback'].apply(preprocess_text)
reference['processed_refer'] = reference['refer_feedback'].apply(preprocess_text)

# Initialize TF-IDF vectorizer and fit on reference data
tfidf = TfidfVectorizer()
reference_tfidf = tfidf.fit_transform(reference['processed_refer'])

# Match each feedback entry to the most similar reference
matched_labels = []
for feedback in data_feed['processed_feedback']:
    # Convert current feedback to TF-IDF vector
    feedback_vec = tfidf.transform([feedback])
    # Calculate similarity scores
    similarities = cosine_similarity(feedback_vec, reference_tfidf)[0]
    # Get index of most similar reference entry
    top_match_idx = similarities.argmax()
    # Pull the corresponding labels
    matched_label = reference.iloc[top_match_idx][['sub-category', 'category', 'sentiment']]
    matched_labels.append(matched_label)

# Merge labels back into data_feed
data_feed = pd.concat([data_feed, pd.DataFrame(matched_labels).reset_index(drop=True)], axis=1)
# Clean up the temporary processed column (optional)
data_feed = data_feed.drop('processed_feedback', axis=1)

# View the result
print(data_feed)

Key Notes

  • Preprocessing helps reduce noise in text (like punctuation or case differences) that could throw off similarity calculations
  • Cosine similarity measures how "close" two text vectors are, with scores ranging from 0 (no overlap) to 1 (identical)

Approach 2: Pre-trained Language Model (BERT) for Semantic Matching

If you need more accurate matching (especially for text with different phrasing but the same meaning), use a pre-trained language model like BERT. This method understands semantic context, so it'll correctly pair "Fast Delivery" with "The delivery was on time" even if the wording isn't identical.

Step-by-Step Implementation

We'll use the sentence-transformers library to generate sentence embeddings (numerical representations of text meaning), then compute cosine similarity.

import pandas as pd
from sentence_transformers import SentenceTransformer, util

# Load your datasets (same as above)
data_feed = pd.DataFrame({
    'feedback': [
        'Fast Delivery. Always before time.Thanks',
        'I have order brown shoe .And I got olive green shoe',
        'Delivery guy is a decent nd friendly guy',
        'Its really good .. my daughter loves it',
        'One t shirt was fully crushed rest everything is good',
        'Superfast delivery! I\'m impressed.'
    ]
})

reference = pd.DataFrame({
    'refer_feedback': [
        'The delivery was on time.',
        'he was polite enough',
        'worst products'
    ],
    'sub-category': ['delivery speed', 'delivery man behaviour', 'product quality'],
    'category': ['delivery', 'delivery', 'general'],
    'sentiment': ['positive', 'positive', 'negative']
})

# Load a lightweight, efficient pre-trained model
model = SentenceTransformer('all-MiniLM-L6-v2')

# Generate embeddings for all reference feedbacks
reference_embeddings = model.encode(reference['refer_feedback'], convert_to_tensor=True)

# Match each feedback entry to the most similar reference
matched_labels = []
for feedback in data_feed['feedback']:
    # Generate embedding for current feedback
    feedback_embedding = model.encode(feedback, convert_to_tensor=True)
    # Calculate cosine similarity scores
    cos_scores = util.cos_sim(feedback_embedding, reference_embeddings)[0]
    # Get index of top matching reference
    top_match_idx = cos_scores.argmax().item()
    # Pull the labels
    matched_label = reference.iloc[top_match_idx][['sub-category', 'category', 'sentiment']]
    matched_labels.append(matched_label)

# Merge labels into data_feed
data_feed = pd.concat([data_feed, pd.DataFrame(matched_labels).reset_index(drop=True)], axis=1)

# View the result
print(data_feed)

Key Notes

  • The all-MiniLM-L6-v2 model is fast and lightweight, perfect for this use case
  • This method handles semantic similarity better than TF-IDF—great for cases where feedback uses different words to mean the same thing

Pro Tips for Better Results

  • Add a similarity threshold: Only assign a label if the top similarity score is above a certain threshold (e.g., 0.5). If not, mark the entry as unknown to avoid incorrect matches.
  • Expand your reference dataset: The more diverse your reference feedbacks are, the better the model will handle edge cases in your data_feed.
  • Tweak preprocessing: Depending on your text, you might want to add spell correction, remove special characters, or filter out stopwords to improve matching accuracy.

内容的提问来源于stack exchange,提问作者James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:33:11