You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python字典转为符合Rasa规范的带多行字符串的YAML文件?

问题描述

我想把下面结构的Python字典:

{
    "version": "3.1",
    "nlu": [
        {
            "intent": "greet",
            "examples": ["hi", "hello", "howdy"]
        },
        {
            "intent": "goodbye",
            "examples": ["goodbye", "bye", "see you later"]
        }
    ]
 }

转换成符合Rasa训练数据格式的YAML文件,要求每个examples键值前带竖线|,格式如下:

version: "3.1"
nlu:
- intent: greet
  examples: |
    - hi
    - hello
    - howdy
- intent: goodbye
  examples: |
    - goodbye
    - bye
    - see you later

我尝试过将examples的值转为字符串后用yaml.dump()生成YAML,但可读性极差,示例如下:

version: '3.1'
nlu:
- intent: greet
  examples: "  - hi\n  - hello\n  - howdy\n" 
- intent: goodbye
  examples: "  - goodbye\n  - bye\n  - see you later\n"  

请问有没有最直接的方法实现目标,同时保证生成的YAML可读性?

解决方法

方案1:全局自定义列表序列化

通过PyYAML的自定义representer,让所有列表类型以带|的多行格式输出,刚好匹配examples的需求:

import yaml

def represent_examples_list(dumper, data):
    # 将列表元素转为每行带"- "的缩进格式字符串
    examples_content = "\n".join(f"    - {item}" for item in data)
    # 指定用|标记多行字符串
    return dumper.represent_scalar('tag:yaml.org,2002:str', examples_content, style='|')

# 注册自定义的列表序列化逻辑
yaml.add_representer(list, represent_examples_list)

# 原始数据字典
raw_data = {
    "version": "3.1",
    "nlu": [
        {"intent": "greet", "examples": ["hi", "hello", "howdy"]},
        {"intent": "goodbye", "examples": ["goodbye", "bye", "see you later"]}
    ]
}

# 写入YAML文件
with open('nlu_training_data.yml', 'w', encoding='utf-8') as file:
    yaml.dump(raw_data, file, sort_keys=False, allow_unicode=True)

方案2:仅针对examples字段处理

如果不想全局修改列表的序列化行为,可以只处理包含examples键的字典:

import yaml
from yaml.scalarstring import LiteralScalarString

def represent_intent_dict(dumper, data):
    if isinstance(data, dict) and 'examples' in data:
        # 复制原字典避免修改原始数据
        processed_data = data.copy()
        # 将examples列表转为带格式的多行字符串
        examples_str = "\n".join(f"    - {item}" for item in processed_data['examples'])
        processed_data['examples'] = LiteralScalarString(examples_str)
        return dumper.represent_dict(processed_data)
    # 其他字典按默认逻辑处理
    return dumper.represent_dict(data)

# 注册自定义字典序列化逻辑
yaml.add_representer(dict, represent_intent_dict)

# 后续写入文件代码同方案1

关键说明

  • sort_keys=False:保证YAML中键的顺序和原字典一致,避免字段顺序混乱。
  • LiteralScalarString:直接生成带|的多行字符串格式,比手动拼接更贴合PyYAML的内部逻辑。

内容的提问来源于stack exchange,提问作者vonbecker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 18:11:10