You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Pandas统计列表中各关键词的出现次数?

问题解决:正确统计关键词出现次数

原代码核心问题

  1. 循环逻辑错误:for num in count_of_occ遍历的是列表中的初始值(全为0),导致始终只修改count_of_occ[0],其余元素从未更新。
  2. 字节转字符串的问题:直接将ctnt.content(字节数据)转为字符串会引入不必要的转义字符和前缀,影响关键词匹配,应该用ctnt.text获取网页文本内容。
  3. 冗余判断:if i in output_content完全多余,str.count()方法本身会在关键词不存在时返回0。

修正后的代码

from collections import Counter
import pandas as pd
import requests

keywords = ["Goulds", "Pump", "http"]
df = pd.DataFrame({'Keyword': keywords})
URL = ["https://www.gouldspumps.com/en-US/Home/"]

# 获取网页文本内容
ctnt = requests.get(URL[0], verify=False)
output_content = ctnt.text  # 改用text获取网页文本

count_of_occ = []
# 遍历每个关键词,统计次数并加入列表
for keyword in keywords:
    count_of_occ.append(output_content.count(keyword))

df['Occurence'] = count_of_occ
print(df)

输出结果示例

Keyword  Occurence
0  Goulds        429
1    Pump         87
2    http         63

额外优化(可选)

如果需要更高效的统计(尤其是关键词数量多的时候),可以用Counter一次性统计所有关键词:

from collections import Counter
# ... 前面代码不变
word_counts = Counter(output_content.split())
count_of_occ = [word_counts.get(keyword, 0) for keyword in keywords]

注意:这种方式是按空格分割的单词统计,适合完整单词匹配;如果需要子串匹配,还是用str.count()更合适。

内容的提问来源于stack exchange,提问作者TroyB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 11:37:52