You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从Kaggle CSV生成指定词汇在标题和描述列的出现统计表

解决Python从Netflix数据集中统计指定词汇出现情况的问题

原代码存在的问题

  • 未定义变量words,导致lambda x:len(set(words) & set(x))报错
  • royaltitles行逻辑错误,错误地对字符串'title'做拆分统计,完全偏离需求
  • 未处理description列的词汇统计
  • 缺少生成目标表格的核心逻辑

修正后的完整代码

需求1:生成每个节目(行)的标题、描述及目标词汇出现次数表格

import pandas as pd

# 读取数据集(替换成你的本地文件路径)
df = pd.read_csv(r'\PATH\titles.csv') 

# 指定要统计的皇家相关词汇列表
royal_words = {'crown', 'queen', 'king', 'prince', 'princess', 'royal', 'majesty', 'monarchy', 'sovereignty', 'throne', 'command', 'kingdom', 'duke', 'duchess'}

# 提取需要的列,使用copy避免SettingWithCopyWarning
small_df = df[['title','description']].copy()

# 定义统计函数:计算文本中目标词汇的出现次数
def count_royal_words(text):
    if pd.isna(text):  # 处理空值,避免报错
        return 0
    words = text.lower().split()
    return len([word for word in words if word in royal_words])

# 分别统计标题和描述中的目标词汇数量
small_df['title_royal_count'] = small_df['title'].apply(count_royal_words)
small_df['description_royal_count'] = small_df['description'].apply(count_royal_words)
small_df['total_royal_count'] = small_df['title_royal_count'] + small_df['description_royal_count']

# 生成目标表格,可选筛选出有目标词汇的条目
royaltitles_table = small_df[small_df['total_royal_count'] > 0]

print(royaltitles_table)

需求2:生成每个目标词汇在标题和描述列中的总出现次数表格

import pandas as pd

df = pd.read_csv(r'\PATH\titles.csv') 
royal_words = ['crown', 'queen', 'king', 'prince', 'princess', 'royal', 'majesty', 'monarchy', 'sovereignty', 'throne', 'command', 'kingdom', 'duke', 'duchess']

# 提取所有非空的标题和描述文本,转为小写后合并拆分
all_title_words = ' '.join(df['title'].dropna().astype(str).str.lower()).split()
all_desc_words = ' '.join(df['description'].dropna().astype(str).str.lower()).split()

# 统计每个目标词汇在对应列的出现次数,缺失词汇填充0
title_counts = pd.Series(all_title_words).value_counts().reindex(royal_words, fill_value=0)
desc_counts = pd.Series(all_desc_words).value_counts().reindex(royal_words, fill_value=0)

# 组合成汇总表格
royaltitles_table = pd.DataFrame({
    '标题出现次数': title_counts,
    '描述出现次数': desc_counts,
    '总出现次数': title_counts + desc_counts
})

print(royaltitles_table)

代码说明

  • 用set存储目标词汇(需求1)可提升查找效率,避免重复匹配
  • 加入空值处理逻辑,防止因缺失文本导致报错
  • 需求1的表格聚焦单个节目,展示目标词汇在标题、描述的分布及总数,可快速定位相关内容
  • 需求2的表格聚焦词汇本身,统计全局出现次数,适合分析词汇热度

内容的提问来源于stack exchange,提问作者JFens007

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 04:40:42