You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pydantic进行JSON格式化输出:二维数组的美观打印问题

Pydantic进行JSON格式化输出:二维数组的美观打印问题

嘿,我明白你现在遇到的问题了——你想让二维数组里的每个子数组(也就是你定义的TwoDim类型)在JSON输出里紧凑地显示在一行,而不是被拆分成多行,但目前用自定义编码器得到的是带引号的字符串形式,这显然不是你想要的合法JSON结构对吧?

咱们先来分析下你当前代码的问题:你的NoIndentEncoder.encode方法返回的是一个字符串(比如"[0.0, 1.2]"),当Pydantic把这个结果传给JSON编码器的时候,它会把这个字符串当成普通的字符串值,所以自动加上了引号,导致输出里的子数组变成了字符串类型,而不是真正的JSON数组。

接下来给你两个可行的解决方案:


方案一:自定义JSON dumps函数,精准控制格式化

这个方法是直接自定义整个JSON序列化的过程,遍历你的数据结构,把TwoDim类型的子数组标记出来,然后在格式化的时候让它们紧凑显示,其他部分保持4空格的缩进。

代码示例如下:

import json
from typing import Sequence
from annotated_types import Len
from typing_extensions import Annotated
from pydantic import BaseModel, Field, ConfigDict

TwoDim = Annotated[
    Sequence[float],
    Len(min_length=2, max_length=2),
]

class Tester(BaseModel):
    test_val: Sequence[TwoDim] = Field()

def custom_json_dumps(obj, indent=4):
    # 先把Pydantic模型转成Python原生类型(字典/列表)
    if isinstance(obj, BaseModel):
        obj = obj.model_dump()
    
    def format_item(item):
        # 判断是否是符合TwoDim定义的数组(长度为2的float序列)
        if isinstance(item, list) and len(item) == 2 and all(isinstance(x, float) for x in item):
            return json.dumps(item)  # 紧凑序列化这个子数组
        elif isinstance(item, list):
            return [format_item(i) for i in item]
        elif isinstance(item, dict):
            return {k: format_item(v) for k, v in item}
        else:
            return item
    
    # 先处理所有TwoDim子数组,得到带有字符串形式子数组的结构
    processed_obj = format_item(obj)
    
    # 定义自定义编码器,把字符串形式的子数组解析回原生数组
    class CustomEncoder(json.JSONEncoder):
        def encode(self, obj):
            if isinstance(obj, list):
                new_list = []
                for elem in obj:
                    if isinstance(elem, str) and elem.startswith('[') and elem.endswith(']'):
                        try:
                            new_list.append(json.loads(elem))
                        except:
                            new_list.append(elem)
                    else:
                        new_list.append(elem)
                return super().encode(new_list)
            return super().encode(obj)
    
    # 最后用自定义编码器和指定缩进输出
    return json.dumps(processed_obj, indent=indent, cls=CustomEncoder)

# 测试代码
test_data = Tester(test_val=[[0.0, 1.2], [3.0, 1.4], [5.6, 7.8]])
# 输出到文件
with open('output.json', 'w') as f:
    f.write(custom_json_dumps(test_data))

这个方案会输出你想要的结构:

{
    "test_val": [
        [0.0, 1.2],
        [3.0, 1.4],
        [5.6, 7.8]
    ]
}

方案二:改进自定义编码器,避免字符串转义

另一种思路是改进你的自定义编码器,让它直接处理整个JSON生成过程,而不是只针对TwoDim类型。我们可以利用JSONEncoder的default方法来处理特定类型,而不是重载encode方法:

import json
from typing import Sequence
from annotated_types import Len
from typing_extensions import Annotated
from pydantic import BaseModel, Field, ConfigDict

TwoDim = Annotated[
    Sequence[float],
    Len(min_length=2, max_length=2),
]

class NoIndentEncoder(json.JSONEncoder):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self._indent = kwargs.get('indent', 0)
        self._current_indent = 0

    def encode(self, obj):
        # 判断是否是TwoDim类型的序列
        if isinstance(obj, list) and all(isinstance(item, list) and len(item)==2 and all(isinstance(x, float) for x in item) for item in obj):
            lines = []
            self._current_indent += self._indent
            indent_str = ' ' * self._current_indent
            for item in obj:
                lines.append(f'{indent_str}{json.dumps(item)}')
            self._current_indent -= self._indent
            return f'[\n{",\n".join(lines)}\n{" " * self._current_indent}]'
        return super().encode(obj)

class Tester(BaseModel):
    test_val: Sequence[TwoDim] = Field()
    model_config = ConfigDict(
        json_encoders={
            Sequence[TwoDim]: NoIndentEncoder().encode
        }
    )

# 测试代码
test_data = Tester(test_val=[[0.0, 1.2], [3.0, 1.4], [5.6, 7.8]])
# 输出到文件
with open('output.json', 'w') as f:
    f.write(test_data.model_dump_json(indent=4))

这个方案同样能生成符合要求的JSON结构,而且直接利用了Pydantic的json_encoders配置,更贴合你的初始思路。


总结一下,核心问题就是你之前的编码器返回了字符串,导致被JSON再次转义。只要调整编码器的实现,让它直接生成正确的JSON数组结构,而不是字符串,就能解决这个问题啦。

备注:内容来源于stack exchange,提问作者nottelling

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.16 08:52:59