Python pandas将任意深度嵌套字典列表转两列问答DataFrame
适配任意嵌套深度的QnA结果展平方案
问题场景
- 开发新闻类问答解析NLP模型时,受最小运行RAM约束,采用逐句拆分解析、结果逐次追加的处理逻辑
- 输出结果嵌套结构无固定规则,嵌套深度随输入文章的总篇数、段落数、句子数随机变化
- 直接调用
pd.DataFrame(data)加载原始数据会得到13578行×32列的高维表,内部残留大量嵌套结构,使用flatten-dict、常规深度展平工具均无法得到预期结果,手动指定展平字段会固定触发报错 - 目标输出:仅包含
question、answer两个字段的标准结构化DataFrame
原始数据结构示例
[ [ {'answer': 'Meta Platforms ', 'question': 'What is the name of the company that became two of the most talked-about social media companies in recent months?'}, {'answer': ' Twitter ', 'question': 'What was the name of the two most talked-about social media companies in recent months?'}, {'answer': 'Elon Musk ', 'question': 'Who took a 9.2% stake in Twitter?'} ], [ {'answer': '$54.20 ', 'question': 'How much did Musk bid to acquire Twitter?'}, {'answer': ' $44 billion ', 'question': 'How much did Musk bid to acquire Twitter?'} ], # 后续多层嵌套结构省略,深度随机 ]
实现代码
核心逻辑用递归生成器遍历所有嵌套层级,不需要提前预判嵌套深度,内存占用极低:
import pandas as pd from collections.abc import Iterable def iter_qna(nested_input): # 遍历可迭代对象,跳过字典、字符串避免拆分层级/拆分单个字符 if isinstance(nested_input, Iterable) and not isinstance(nested_input, (dict, str, bytes)): for element in nested_input: yield from iter_qna(element) # 命中标准问答字典时提取字段,自动清除首尾多余空格 elif isinstance(nested_input, dict) and {"question", "answer"}.issubset(nested_input.keys()): yield { "question": nested_input["question"].strip(), "answer": nested_input["answer"].strip() } # 替换成你自己存储原始解析结果的变量名即可 result_df = pd.DataFrame(iter_qna(raw_data))
方案特性
- 适配任意嵌套深度:无论外层包裹多少层列表、元组,只要叶子节点是包含
question和answer键的字典,就能正确提取,不需要提前感知数据结构 - 低内存占用:采用生成器逐元素遍历,不会一次性加载全量展平后的中间数据,符合最小RAM的运行约束
- 容错性强:解析过程中产生的乱码、广告注入条目(比如示例中的
⁇!-- googletag.cmd.push类异常内容)只要字典结构符合规则就会正常保留,结构异常的条目自动跳过,不会触发全局报错 - 自带基础清洗:提取字段时自动清除问答文本首尾的多余空格,解决示例中字段值前后带冗余空格的问题
内容的提问来源于stack exchange,提问作者NaNFill
相关产品推荐
相关产品推荐

