You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将Polars DataFrame转换为指定嵌套JSON格式的实现咨询

将Polars DataFrame转换为指定嵌套JSON格式的实现咨询

我有一个包含产品名称、问题和答案的DataFrame,想要把它处理转换成指定的JSON格式,每个产品需要嵌套对应的问题和答案模块。

我的DataFrame代码如下:

import polars as pl

df = pl.DataFrame({
    "Product": ["X", "X", "Y", "Y"],
    "Question": ["Q1", "Q2", "Q3", "Q4"],
    "Anwers": ["A1", "A2", "A3", "A4"],
}) 

期望的JSON输出:

{
    "faqByCommunity": {
        "id": 5,
        "communityName": "name",
        "faqList": [
            {
                "id": 1,
                "product": "X",
                "faqs": [
                    {
                        "id": 1,
                        "question": "Q1",
                        "answer": "A1"
                    },
                    {
                        "id": 2,
                        "question": "Q2",
                        "answer": "A2"
                    }

                ]
            },
            {
                "id": 2,
                "product": "Y",
                "faqs": [
                    {
                        "id": 1,
                        "question": "Q3",
                        "answer": "A3"
                    },
                    {
                        "id": 2,
                        "question": "Q4",
                        "answer": "A4"
                    }

                ]
            }
        ]
    }
}

因为输出的外层部分是固定的,我想可以在Polars生成内容前后把静态部分拼接进去,但现在不确定怎么处理里面的嵌套部分。


没问题,咱们可以分两步来实现:先处理Polars DataFrame生成嵌套的产品FAQ结构,再把静态的外层结构套上去,全程用Polars原生功能+Python标准库就能搞定。

第一步:用Polars生成嵌套的产品FAQ列表

核心思路是先给每个问题加ID,再按产品分组嵌套FAQ集合,最后给每个产品条目加ID:

import polars as pl
import json

# 初始化你的DataFrame
df = pl.DataFrame({
    "Product": ["X", "X", "Y", "Y"],
    "Question": ["Q1", "Q2", "Q3", "Q4"],
    "Anwers": ["A1", "A2", "A3", "A4"],
}) 

# 1. 给每个问题添加ID(按顺序从1开始),同时修正列名拼写错误*Anwers*→answer
df_with_ids = df.with_row_index(name="id", offset=1).rename({"Anwers": "answer", "Question": "question"})

# 2. 按产品分组,把每个产品下的问题+答案嵌套成faqs列表,再给每个产品条目加ID
nested_faqs = (
    df_with_ids
    .group_by("Product", maintain_order=True)
    .agg(pl.struct(["id", "question", "answer"]).alias("faqs"))
    .with_row_index(name="id", offset=1)
    .rename({"Product": "product"})
)

# 3. 把Polars的结构化数据转成Python字典列表,方便后续拼接
faq_list = nested_faqs.to_dicts()

第二步:拼接静态外层并生成最终JSON

现在把固定的外层结构和动态生成的faq_list组合,再转成格式化的JSON:

# 定义固定的外层结构
final_result = {
    "faqByCommunity": {
        "id": 5,
        "communityName": "name",
        "faqList": faq_list
    }
}

# 转成带缩进的美观JSON字符串
formatted_json = json.dumps(final_result, indent=4)
print(formatted_json)

运行这段代码后,就能得到和你期望完全一致的JSON输出了。这里的优势是全程用Polars的原生分组嵌套功能,比手动循环拼接字典高效得多,尤其是处理大数据量的时候;同时with_row_index自动生成ID,避免了手动计数的麻烦。

备注:内容来源于stack exchange,提问作者Simon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 18:39:36