You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计DataFrame关键词在指定文本中的出现频率并新增列

解决方案

实现思路

  • 统一文本大小写并去除标点,避免大小写、标点对统计结果的干扰
  • 统计文本内各单词的出现次数
  • 将统计结果与原DataFrame关联,未在文本中出现的单词频率填充为0

代码实现

import pandas as pd
import re

# 给定的文本与初始DataFrame
text = """Alice has two apples and bananas. Apples are very healty."""
df = pd.DataFrame({'word': ['apples', 'bananas', 'company']})

# 预处理文本:转小写、移除标点、拆分单词
processed_text = re.sub(r'[^\w\s]', '', text.lower())
word_list = processed_text.split()

# 统计单词出现频率
word_count = pd.Series(word_list).value_counts().reset_index()
word_count.columns = ['word', 'frequency']

# 合并数据并填充缺失值
result_df = df.merge(word_count, on='word', how='left').fillna(0)
result_df['frequency'] = result_df['frequency'].astype(int)

print(result_df)

最终输出

wordfrequency
apples2
bananas1
company0

内容的提问来源于stack exchange,提问作者Kas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 18:25:23