You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用spaCy从DataFrame列生成分词列表并统计词频?

问题2:用Counter统计全量词频

Counter.update()是原地修改Counter对象,没有返回值,所以用map或apply调用时,会返回由None组成的Series,但Counter已经被更新。你需要直接输出Counter对象,而非map的返回结果:

from collections import Counter

counter_df = Counter()

# 遍历所有分词列表,更新词频
for token_list in example_df["tokens"]:
    counter_df.update(token_list)

# 输出统计结果
print(counter_df)

或者用apply(虽然会返回None的Series,但Counter已完成更新):

example_df["tokens"].apply(counter_df.update)
print(counter_df)

执行后会得到预期的词频统计:

Counter({'a': 4, 'All': 2, 'roads': 1, 'lead': 1, 'to': 1, 'Rome': 1, 'work': 1, 'and': 1, 'no': 1, 'play': 1, 'makes': 1, 'Jack': 1, 'dull': 1, 'buy': 1, 'Any': 1, 'port': 1, 'in': 1, 'storm': 1, 'Avoid': 1, 'questioner': 1, ',': 1, 'for': 1, 'he': 1, 'is': 1, 'also': 1, 'tattler': 1})

内容的提问来源于stack exchange,提问作者Ahmad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 17:18:24