You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Polars从迭代器生成词频表并将词频字典转为Polars DataFrame

如何用Polars从迭代器生成词频表并将词频字典转为Polars DataFrame

嘿,我来帮你搞定这个Polars词频统计的问题!其实Polars本身就有高效的内置方法,不管是直接从单词列表生成词频表,还是把现成的词频字典转成DataFrame,都很简单,我分两种情况给你说明:

一、直接从单词列表生成词频DataFrame(无需字典)

你之前用defaultdict循环统计的方式虽然可行,但Polars的矢量化操作效率更高,尤其是处理大数据量的时候。直接用Polars的分组聚合就能一步到位:

import polars as pl

# 示例单词列表
list_of_words = ["apple", "banana", "apple", "orange", "banana", "apple"]

# 生成词频DataFrame
word_freq_df = (
    pl.Series(list_of_words, name="word")
    .to_frame()  # 转成DataFrame
    .group_by("word")  # 按单词分组
    .agg(pl.count().alias("count"))  # 统计每组的数量并命名为count
    .sort("count", descending=True)  # 可选:按词频从高到低排序,方便查看
)

print(word_freq_df)

运行后会得到这样的结果:

shape: (3, 2)
┌────────┬───────┐
│ word   ┆ count │
│ ---    ┆ ---   │
│ str    ┆ u32   │
╞════════╪═══════╡
│ apple  ┆ 3     │
│ banana ┆ 2     │
│ orange ┆ 1     │
└────────┴───────┘

这种方式完全跳过了字典的中间步骤,代码更简洁,性能也更好。

二、将已有的词频字典转为Polars DataFrame

如果已经通过其他方式得到了词频字典,也有两种简单的转换方法:

方法1:直接从字典的items构建

字典的items()方法会返回(key, value)的元组列表,直接传给pl.DataFrame并指定列名即可:

from collections import defaultdict

# 先得到你的词频字典
word_freq = defaultdict(int)
for word in list_of_words:
    word_freq[word] += 1

# 转换为DataFrame
df_from_dict = pl.DataFrame(word_freq.items(), schema=["word", "count"])

方法2:使用pl.from_dict方法

这个方法需要把字典调整为列优先的结构(即键是列名,值是列的内容列表):

df_from_dict2 = pl.from_dict({
    "word": list(word_freq.keys()),
    "count": list(word_freq.values())
})

两种方法都能得到相同的结果,选哪种全看你的习惯——第一种更简洁,第二种适合你想明确控制每一列内容的场景。

备注:内容来源于stack exchange,提问作者ste_kwr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 09:18:09