You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python数据迭代合并:关联输入字段、原始答案与人工答案

匹配表格列、原始答案与人工答案并生成指定格式输出

问题背景

我是Python新手,刚开始做数据处理工作,需要合并不同数据对象,生成易读的对比展示内容。

处理的数据

{"flowDefinitionArn": "arn:aws:sagemaker:us-east-1:2345:flow-definition/definition_name",
"humanAnswers": [
    {
        "acceptanceTime": "2022-11-15T18:37:50.085Z",
        "answerContent": {
            "extracted1_1": "Italy",
            "extracted1_2": "Rome",
            "extracted1_3": "5555",
            "extracted2_1": "Czech",
            "extracted2_2": "Prague",
            "extracted2_3": "3333",
            "reportDate": "2022-06-01T08:30",
            "reportOwner": "John Smith"
        },
        "submissionTime": "2022-11-15T18:38:32.791Z",
        "timeSpentInSeconds": 42.706,
        "workerId": "1234",
        "workerMetadata": {
            "identityData": {
                "identityProviderType": "Cognito",
                "issuer": "https://cognito-idp.us-east-1.amazonaws.com/",
                "sub": "c"
            }
        }
    }
],
"humanLoopName": "test",
"inputContent": {
    "document": {
        "documentType": "countryReport",
        "fields": [
            {
                "id": "reportOwner",
                "type": "string",
                "validation": "",
                "value": "John Smith"
            },
            {
                "id": "reportDate",
                "type": "date",
                "validation": "",
                "value": "2022-06-01T08:30"
            },
            {
                "id": "locationList",
                "type": "table",
                "value": {
                    "columns": [
                        {
                            "id": "country",
                            "type": "string"
                        },
                        {
                            "id": "capital",
                            "type": "string"
                        },
                        {
                            "id": "population",
                            "type": "number"
                        }
                    ],
                    "rows": [
                        [
                            "UK",
                            "London",
                            1234
                        ],
                        [
                            "France",
                            "Paris",
                            321
                        ]
                    ]
                }
            }
        ]
    },
    "document_types": [
        {
            "displayName": "Email",
            "id": "email"
        },
        {
            "displayName": "Invoice",
            "id": "invoice"
        },
        {
            "displayName": "Other",
            "id": "other"
        }
    ],
    "input_s3_uri": "s3://my-input-bucket/file1.pdf"
}
}

期望输出格式

Input info: country, Original answer: UK, Human answer: extracted1_1: Italy

Input info: capital, Original answer: London, Human answer: extracted1_2: Rome

Input info: population, Original answer: 1234, Human answer: extracted1_3: 5555

Input info: country, Original answer: France, Human answer: extracted2_1: Czech

Input info: capital, Original answer: Paris, Human answer: extracted2_2: Prague

Input info: population, Original answer: 321, Human answer: extracted2_3: 3333

当前代码

s3_client       = boto3.client('s3')
response        = s3_client.get_object(Bucket=f'{config["bucket"]}', Key=f'{config["file_name"]}')
data            = response['Body'].read()
d               = json.loads(data)
column          = d['inputContent']['document']['fields'][2]['value']['columns']
row             = d['inputContent']['document']['fields'][2]['value']['rows']
answers         = d['humanAnswers'][0]['answerContent']
str_row         = str(row)
iter_col        = iter(column)
iter_row        = iter(str_row)
combined        = ''

for a in answers.items():
    nxt_col = next(iter_col)
    for list in row:
        for values in list:
            v = values
            combined += str(v + ", ")


print(f'Input info: {nxt_col}, Original Answer: {str_row}, Human Answer: {a}')

解决方案

首先要理清数据对应逻辑:原始表格的2行数据,分别对应人工答案里的extracted1_*和extracted2_*组;每组里的_1/_2/_3依次对应表格的country/capital/population列。

修改后的代码如下:

import boto3
import json

# 读取S3数据(假设config已定义)
s3_client = boto3.client('s3')
response = s3_client.get_object(Bucket=f'{config["bucket"]}', Key=f'{config["file_name"]}')
data = response['Body'].read()
d = json.loads(data)

# 提取关键数据
columns = d['inputContent']['document']['fields'][2]['value']['columns']
original_rows = d['inputContent']['document']['fields'][2]['value']['rows']
human_answers = d['humanAnswers'][0]['answerContent']

# 整理列名列表
col_names = [col['id'] for col in columns]

# 过滤并分组人工答案:只保留extracted开头的,按行号分组
extracted_answers = {}
for key, value in human_answers.items():
    if key.startswith('extracted'):
        row_num = key.split('_')[1]
        if row_num not in extracted_answers:
            extracted_answers[row_num] = []
        extracted_answers[row_num].append((key, value))

# 按行号排序,确保顺序正确
sorted_rows = sorted(extracted_answers.keys())

# 遍历每一行数据,生成对比内容
for idx, row_num in enumerate(sorted_rows):
    original_row = original_rows[idx]
    human_row = extracted_answers[row_num]
    # 遍历每一列的对应数据
    for col_idx, (col_name, original_val) in enumerate(zip(col_names, original_row)):
        human_key, human_val = human_row[col_idx]
        print(f'Input info: {col_name}, Original answer: {original_val}, Human answer: {human_key}: {human_val}')
    # 每行之间空一行
    print()

代码说明

  1. 提取列名:从表格columns中取出每个列的id,得到['country', 'capital', 'population']
  2. 分组人工答案:将extracted1_*和extracted2_*分别归为两组,对应原始数据的两行
  3. 匹配对应关系:遍历每一行原始数据,同时对应一组人工答案,按列索引匹配列名、原始值和人工答案的键值对
  4. 生成输出:按照期望格式打印每条对比内容,行与行之间空一行

内容的提问来源于stack exchange,提问作者RunRabbit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:05:20