You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于NLP统计数据集A词汇/语句在数据集B的出现频率(含相似匹配及关联City)

解决方案:跨数据集相似文本匹配与统计可视化

一、核心思路梳理

你的需求核心是相似文本匹配+关联City字段统计+可视化呈现,不用单独处理数据集A再关联B,直接以A的目标文本(如'Tom has a daughter'及相关关键词)为基准,和数据集B的每条记录做匹配,同时结合City字段过滤或分组即可。

二、步骤拆解与代码实现

1. 通用文本预处理函数

先写一个统一的预处理函数,消除大小写、标点、停用词的干扰,确保匹配准确性:

import pandas as pd
import nltk
from nltk.corpus import stopwords
from string import punctuation
nltk.download('stopwords')

def preprocess_text(text):
    # 转小写
    text = text.lower()
    # 去除标点符号
    text = ''.join([c for c in text if c not in punctuation])
    # 过滤停用词
    stop_words = set(stopwords.words('english'))
    filtered_words = [word for word in text.split() if word not in stop_words]
    return ' '.join(filtered_words)

2. 加载并预处理数据集

对两个数据集的Name字段做预处理,保留City字段用于后续关联:

# 加载数据集A
data = {'Name': ['Tom has a daughter', 'Joseph likes to fish', 'Krish is a new student/employee', 'John! What are you doing?'], 'City': ['London', 'Bristol', 'Leeds', 'London']}  
df1 = pd.DataFrame(data)
# 加载数据集B
data1 = {'Name': ['Krish is a new student/employee', 'The sky is blue', 'We are all humans', 'Tom has a daughter'], 'City': ['Leeds', 'Bristol', 'Leeds', 'London']}  
df2 = pd.DataFrame(data1)

# 预处理Name字段
df1['processed_name'] = df1['Name'].apply(preprocess_text)
df2['processed_name'] = df2['Name'].apply(preprocess_text)

3. 相似匹配与统计(结合City字段)

提供两种匹配方式,按需选择:

方式1:关键词匹配(简单直接,适配含'Tom'/'daughter'的需求)

从目标语句提取核心关键词,在B中筛选包含任一关键词的记录,同时支持City过滤:

# 从A的目标语句提取核心关键词
target_keywords = {'tom', 'daughter'}

# 判断单条文本是否匹配关键词
def match_keywords(text):
    text_words = set(text.split())
    return len(text_words & target_keywords) > 0

# 基础匹配(仅文本)
df2['is_match'] = df2['processed_name'].apply(match_keywords)
basic_count = df2['is_match'].sum()
print(f"仅文本匹配频次:{basic_count}")

# 结合City过滤(例:仅统计London地区的匹配)
london_match_count = df2[(df2['is_match']) & (df2['City'] == 'London')]['is_match'].sum()
print(f"London地区匹配频次:{london_match_count}")

方式2:余弦相似度匹配(针对语句整体相似性)

如果需要匹配和'Tom has a daughter'整体语义相似的语句,用TF-IDF计算余弦相似度:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# 获取A中目标语句的预处理结果
target_text = df1[df1['Name'] == 'Tom has a daughter']['processed_name'].iloc[0]
# 构建TF-IDF矩阵
corpus = [target_text] + df2['processed_name'].tolist()
vectorizer = TfidfVectorizer()
tfidf_matrix = vectorizer.fit_transform(corpus)

# 计算目标语句与B中每条语句的相似度
similarities = cosine_similarity(tfidf_matrix[0:1], tfidf_matrix[1:])[0]
# 设置相似度阈值(可按需调整,0.5为中等匹配度)
threshold = 0.5
df2['similarity_match'] = similarities > threshold

similar_count = df2['similarity_match'].sum()
print(f"余弦相似度匹配(阈值{threshold})频次:{similar_count}")

# 结合City过滤
london_similar_count = df2[(df2['similarity_match']) & (df2['City'] == 'London')]['similarity_match'].sum()
print(f"London地区相似度匹配频次:{london_similar_count}")

4. 可视化结果

用柱状图按City分组展示匹配频次,直观呈现分布:

import matplotlib.pyplot as plt
import seaborn as sns

# 按City分组统计关键词匹配数
match_by_city = df2[df2['is_match']].groupby('City').size().reset_index(name='count')
# 补全无匹配的城市,保证图表完整性
all_cities = df2['City'].unique()
match_by_city = match_by_city.set_index('City').reindex(all_cities, fill_value=0).reset_index()

# 绘制柱状图
plt.figure(figsize=(8,5))
sns.barplot(x='City', y='count', data=match_by_city)
plt.title('各城市匹配语句频次(关键词匹配)')
plt.xlabel('城市')
plt.ylabel('匹配次数')
plt.show()

# 若用相似度匹配,将上述代码中的'is_match'替换为'similarity_match'即可

三、关键注意事项

  • 预处理一致性:A和B的文本必须用同一个预处理函数,否则会导致匹配偏差
  • 阈值灵活调整:余弦相似度的阈值可根据需求修改,严格匹配设0.7+,宽松匹配设0.3-
  • City字段拓展:除了过滤,还可以按City做交叉统计(如不同城市的匹配占比),适配更多分析场景

内容的提问来源于stack exchange,提问作者Zara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 08:20:38