如何将Polars中value_counts的结构体列表转为JSON格式存入DataFrame?
Polars列表词频统计转JSON字典字符串实现方案
需求说明
需要对Polars DataFrame中字符串列表类型的列进行词频统计,并将统计结果以JSON格式的字典字符串存入DataFrame。现有输入DataFrame及当前实现的问题如下:
输入DataFrame代码
import polars as pl pl.Config(fmt_table_cell_list_len=6, fmt_str_lengths=100) df = pl.DataFrame({'a': [['the', 'dog', 'is', 'good', 'a'], ['toto', 'tata', 'I']]})
输入表格输出:
shape: (2, 1) ┌───────────────────────────────────┐ │ a │ │ --- │ │ list[str] │ ╞═══════════════════════════════════╡ │ ["the", "dog", "is", "good", "a"] │ │ ["toto", "tata", "I"] │ └───────────────────────────────────┘
当前实现及问题
执行以下代码后,得到的是结构体列表形式的词频结果:
df.with_columns( pl.col('a').list.eval(pl.element().value_counts()) )
输出结果:
shape: (2, 1) ┌───────────────────────────────────────────────────────┐ │ a │ │ --- │ │ list[struct[2]] │ ╞═══════════════════════════════════════════════════════╡ │ [{"is",1}, {"the",1}, {"dog",1}, {"good",1}, {"a",1}] │ │ [{"tata",1}, {"I",1}, {"toto",1}] │ └───────────────────────────────────────────────────────┘
需要将上述结果转换为JSON格式的字典字符串,期望输出如下:
expected = pl.DataFrame({ "a": [ '{"the": 1, "dog": 1, "is": 1, "good": 1, "a": 1}', '{"toto": 1, "tata": 1, "I": 1}' ] })
对应表格输出:
shape: (2, 1) ┌──────────────────────────────────────────────────┐ │ a │ │ --- │ │ str │ ╞══════════════════════════════════════════════════╡ │ {"the": 1, "dog": 1, "is": 1, "good": 1, "a": 1} │ │ {"toto": 1, "tata": 1, "I": 1} │ └──────────────────────────────────────────────────┘
同时需要确认:Polars中value_counts输出的结构体列表能否直接转为JSON格式?
解决方案
方法一:直接转换结构体列表为JSON字符串
利用map_elements遍历结构体列表,将每个结构体转为键值对后序列化:
import polars as pl import json df = pl.DataFrame({'a': [['the', 'dog', 'is', 'good', 'a'], ['toto', 'tata', 'I']]}) # 生成词频统计并转为JSON字典字符串 result = df.with_columns( pl.col('a') .list.value_counts() # 等价于list.eval(pl.element().value_counts()) .map_elements( lambda struct_list: json.dumps({item['value']: item['count'] for item in struct_list}), return_dtype=pl.String ) ) print(result)
方法二:更高效的批量处理方式(避免逐行映射)
如果数据量较大,可通过展开-分组聚合-序列化的方式实现,性能更优:
import polars as pl import json df = pl.DataFrame({'a': [['the', 'dog', 'is', 'good', 'a'], ['toto', 'tata', 'I']]}) result = ( df .with_row_index(name='idx') # 添加行索引用于分组 .explode('a') # 展开列表 .group_by(['idx', 'a']) .agg(count=pl.count()) # 统计词频 .group_by('idx') .agg(pl.struct('a', 'count')) # 重组为结构体列表 .with_columns( pl.col('a') .map_elements(lambda x: json.dumps({d['a']: d['count'] for d in x}), return_dtype=pl.String) ) .drop('idx') ) print(result)
问题解答
- 可以将词频统计结果转换为期望的JSON格式字典字符串:上述两种方法均可实现,最终输出与期望结果一致。
- 结构体列表可以转为JSON格式:Polars的结构体在Python环境中可直接作为类字典对象访问,只需遍历结构体列表生成键值对字典,再通过
json.dumps()序列化即可得到JSON字符串。
内容的提问来源于stack exchange,提问作者dams
相关产品推荐
相关产品推荐

