基于NLP统计数据集A词汇/语句在数据集B的出现频率(含相似匹配及关联City)
解决方案:跨数据集相似文本匹配与统计可视化
一、核心思路梳理
你的需求核心是相似文本匹配+关联City字段统计+可视化呈现,不用单独处理数据集A再关联B,直接以A的目标文本(如'Tom has a daughter'及相关关键词)为基准,和数据集B的每条记录做匹配,同时结合City字段过滤或分组即可。
二、步骤拆解与代码实现
1. 通用文本预处理函数
先写一个统一的预处理函数,消除大小写、标点、停用词的干扰,确保匹配准确性:
import pandas as pd import nltk from nltk.corpus import stopwords from string import punctuation nltk.download('stopwords') def preprocess_text(text): # 转小写 text = text.lower() # 去除标点符号 text = ''.join([c for c in text if c not in punctuation]) # 过滤停用词 stop_words = set(stopwords.words('english')) filtered_words = [word for word in text.split() if word not in stop_words] return ' '.join(filtered_words)
2. 加载并预处理数据集
对两个数据集的Name字段做预处理,保留City字段用于后续关联:
# 加载数据集A data = {'Name': ['Tom has a daughter', 'Joseph likes to fish', 'Krish is a new student/employee', 'John! What are you doing?'], 'City': ['London', 'Bristol', 'Leeds', 'London']} df1 = pd.DataFrame(data) # 加载数据集B data1 = {'Name': ['Krish is a new student/employee', 'The sky is blue', 'We are all humans', 'Tom has a daughter'], 'City': ['Leeds', 'Bristol', 'Leeds', 'London']} df2 = pd.DataFrame(data1) # 预处理Name字段 df1['processed_name'] = df1['Name'].apply(preprocess_text) df2['processed_name'] = df2['Name'].apply(preprocess_text)
3. 相似匹配与统计(结合City字段)
提供两种匹配方式,按需选择:
方式1:关键词匹配(简单直接,适配含'Tom'/'daughter'的需求)
从目标语句提取核心关键词,在B中筛选包含任一关键词的记录,同时支持City过滤:
# 从A的目标语句提取核心关键词 target_keywords = {'tom', 'daughter'} # 判断单条文本是否匹配关键词 def match_keywords(text): text_words = set(text.split()) return len(text_words & target_keywords) > 0 # 基础匹配(仅文本) df2['is_match'] = df2['processed_name'].apply(match_keywords) basic_count = df2['is_match'].sum() print(f"仅文本匹配频次:{basic_count}") # 结合City过滤(例:仅统计London地区的匹配) london_match_count = df2[(df2['is_match']) & (df2['City'] == 'London')]['is_match'].sum() print(f"London地区匹配频次:{london_match_count}")
方式2:余弦相似度匹配(针对语句整体相似性)
如果需要匹配和'Tom has a daughter'整体语义相似的语句,用TF-IDF计算余弦相似度:
from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.metrics.pairwise import cosine_similarity # 获取A中目标语句的预处理结果 target_text = df1[df1['Name'] == 'Tom has a daughter']['processed_name'].iloc[0] # 构建TF-IDF矩阵 corpus = [target_text] + df2['processed_name'].tolist() vectorizer = TfidfVectorizer() tfidf_matrix = vectorizer.fit_transform(corpus) # 计算目标语句与B中每条语句的相似度 similarities = cosine_similarity(tfidf_matrix[0:1], tfidf_matrix[1:])[0] # 设置相似度阈值(可按需调整,0.5为中等匹配度) threshold = 0.5 df2['similarity_match'] = similarities > threshold similar_count = df2['similarity_match'].sum() print(f"余弦相似度匹配(阈值{threshold})频次:{similar_count}") # 结合City过滤 london_similar_count = df2[(df2['similarity_match']) & (df2['City'] == 'London')]['similarity_match'].sum() print(f"London地区相似度匹配频次:{london_similar_count}")
4. 可视化结果
用柱状图按City分组展示匹配频次,直观呈现分布:
import matplotlib.pyplot as plt import seaborn as sns # 按City分组统计关键词匹配数 match_by_city = df2[df2['is_match']].groupby('City').size().reset_index(name='count') # 补全无匹配的城市,保证图表完整性 all_cities = df2['City'].unique() match_by_city = match_by_city.set_index('City').reindex(all_cities, fill_value=0).reset_index() # 绘制柱状图 plt.figure(figsize=(8,5)) sns.barplot(x='City', y='count', data=match_by_city) plt.title('各城市匹配语句频次(关键词匹配)') plt.xlabel('城市') plt.ylabel('匹配次数') plt.show() # 若用相似度匹配,将上述代码中的'is_match'替换为'similarity_match'即可
三、关键注意事项
- 预处理一致性:A和B的文本必须用同一个预处理函数,否则会导致匹配偏差
- 阈值灵活调整:余弦相似度的阈值可根据需求修改,严格匹配设0.7+,宽松匹配设0.3-
- City字段拓展:除了过滤,还可以按City做交叉统计(如不同城市的匹配占比),适配更多分析场景
内容的提问来源于stack exchange,提问作者Zara
相关产品推荐
相关产品推荐

