You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按DataFrame每行生成独立YAML文件的逻辑问题排查

问题原因

你的代码核心错误是循环逻辑嵌套错误:

  1. 外层已经在逐行遍历DataFrame,内层又写了一个遍历全量DataFrame的列表推导式,导致每次循环生成的metrics变量是包含所有行数据的完整列表,不是当前行对应的单条指标
  2. 写入YAML时你把全量列表直接dump进文件,每个文件里存的都是全部指标的内容,自然和文件名对应的单条id数据不匹配,存在大量冗余
  3. 额外的CSV中转逻辑没有必要,还可能因为读CSV时的表头、索引解析问题引入额外的数据错位
  4. 列删除的参数写法不规范,存在pandas版本兼容风险;YAML写入未指定编码,非英文字符容易乱码
修正后可直接运行的代码
import gspread
import pandas as pd
import yaml

# 服务账号授权
gc = gspread.service_account()

# 读取目标Google Sheet工作表
sh = gc.open_by_url('https://docs.google.com/spreadsheets/myspreadhsheet_with_metrics')
worksheet = sh.worksheet('Product/Business metrics v2')

# 直接从工作表数据构建DataFrame,移除无意义的CSV中转步骤
raw_data = worksheet.get_values()
df = pd.DataFrame(raw_data)
# 按原逻辑设置第4行(索引为3)为表头,跳过前面的无效行
df.columns = df.iloc[3]
df = df[4:].reset_index(drop=True)
# 规范drop列写法,显式指定axis参数兼容各版本pandas
df = df.drop('2', axis=1)
# 预处理空值字段
df['Trading impact'] = df['Trading impact'].fillna('').astype(str)

# 逐行生成单条指标对应的YAML文件
for _, row in df.iterrows():
    # 直接构建当前行对应的单条指标字典,不再嵌套全量遍历
    metric = {
        'id': row['Metric'].replace(' ', '_').lower(),
        'type': 'custom',
        'database_id': 'baobab_00894976',
        'schedule': 'hourly',
        'sql': '',
        'display_name': row['Metric'].replace(' ', '_').lower(),
        'definition': row['Official business definition'],
        'description': row['Preliminary simplistic description'],
        'owner': row['Official team ownership (department - tribe)'],
        'team_lead': row['Team lead / Analytics lead'],
        'slack_channel': row['Slack channel'],
        'metric_versions': row['Metric versions'],
        'dimensionality': row['Dimensionality'],
        'data_source': row['Data source (DB)'],
        'looker_available': row['Looker Data model availability'],
        'looker_model': row['Looker Data model name'],
        'source_of_truth': row['Source of truth'],
        'threshold_type': 'gt',
        'threshold_value': row['Target / Threshold'],
        'trading_impact': row['Trading impact'] if row['Trading impact'] != 'nan' else 'not filled'
    }
    # 以指标id为文件名写入,指定utf8编码避免乱码
    with open(f"{metric['id']}.yaml", "w", encoding='utf-8') as f:
        yaml.dump(metric, f, allow_unicode=True, sort_keys=False)
补充说明
  • 代码里加了sort_keys=False参数,生成的YAML会保留你写的字典键顺序,不会自动按字母排序,可读性更好
  • 移除了CSV中转后,少了两次磁盘IO,运行速度更快,也不会出现CSV读写时的索引列、表头错位问题
  • 所有文件写入后,每个YAML文件只会存储对应id的单条指标数据,不会再有冗余内容

内容的提问来源于stack exchange,提问作者Radka Žmers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 06:09:29