如何用Python从Kaggle CSV生成指定词汇在标题和描述列的出现统计表
解决Python从Netflix数据集中统计指定词汇出现情况的问题
原代码存在的问题
- 未定义变量
words,导致lambda x:len(set(words) & set(x))报错 royaltitles行逻辑错误,错误地对字符串'title'做拆分统计,完全偏离需求- 未处理
description列的词汇统计 - 缺少生成目标表格的核心逻辑
修正后的完整代码
需求1:生成每个节目(行)的标题、描述及目标词汇出现次数表格
import pandas as pd # 读取数据集(替换成你的本地文件路径) df = pd.read_csv(r'\PATH\titles.csv') # 指定要统计的皇家相关词汇列表 royal_words = {'crown', 'queen', 'king', 'prince', 'princess', 'royal', 'majesty', 'monarchy', 'sovereignty', 'throne', 'command', 'kingdom', 'duke', 'duchess'} # 提取需要的列,使用copy避免SettingWithCopyWarning small_df = df[['title','description']].copy() # 定义统计函数:计算文本中目标词汇的出现次数 def count_royal_words(text): if pd.isna(text): # 处理空值,避免报错 return 0 words = text.lower().split() return len([word for word in words if word in royal_words]) # 分别统计标题和描述中的目标词汇数量 small_df['title_royal_count'] = small_df['title'].apply(count_royal_words) small_df['description_royal_count'] = small_df['description'].apply(count_royal_words) small_df['total_royal_count'] = small_df['title_royal_count'] + small_df['description_royal_count'] # 生成目标表格,可选筛选出有目标词汇的条目 royaltitles_table = small_df[small_df['total_royal_count'] > 0] print(royaltitles_table)
需求2:生成每个目标词汇在标题和描述列中的总出现次数表格
import pandas as pd df = pd.read_csv(r'\PATH\titles.csv') royal_words = ['crown', 'queen', 'king', 'prince', 'princess', 'royal', 'majesty', 'monarchy', 'sovereignty', 'throne', 'command', 'kingdom', 'duke', 'duchess'] # 提取所有非空的标题和描述文本,转为小写后合并拆分 all_title_words = ' '.join(df['title'].dropna().astype(str).str.lower()).split() all_desc_words = ' '.join(df['description'].dropna().astype(str).str.lower()).split() # 统计每个目标词汇在对应列的出现次数,缺失词汇填充0 title_counts = pd.Series(all_title_words).value_counts().reindex(royal_words, fill_value=0) desc_counts = pd.Series(all_desc_words).value_counts().reindex(royal_words, fill_value=0) # 组合成汇总表格 royaltitles_table = pd.DataFrame({ '标题出现次数': title_counts, '描述出现次数': desc_counts, '总出现次数': title_counts + desc_counts }) print(royaltitles_table)
代码说明
- 用
set存储目标词汇(需求1)可提升查找效率,避免重复匹配 - 加入空值处理逻辑,防止因缺失文本导致报错
- 需求1的表格聚焦单个节目,展示目标词汇在标题、描述的分布及总数,可快速定位相关内容
- 需求2的表格聚焦词汇本身,统计全局出现次数,适合分析词汇热度
内容的提问来源于stack exchange,提问作者JFens007
相关产品推荐
相关产品推荐

