Python数据迭代合并:关联输入字段、原始答案与人工答案
匹配表格列、原始答案与人工答案并生成指定格式输出
问题背景
我是Python新手,刚开始做数据处理工作,需要合并不同数据对象,生成易读的对比展示内容。
处理的数据
{"flowDefinitionArn": "arn:aws:sagemaker:us-east-1:2345:flow-definition/definition_name", "humanAnswers": [ { "acceptanceTime": "2022-11-15T18:37:50.085Z", "answerContent": { "extracted1_1": "Italy", "extracted1_2": "Rome", "extracted1_3": "5555", "extracted2_1": "Czech", "extracted2_2": "Prague", "extracted2_3": "3333", "reportDate": "2022-06-01T08:30", "reportOwner": "John Smith" }, "submissionTime": "2022-11-15T18:38:32.791Z", "timeSpentInSeconds": 42.706, "workerId": "1234", "workerMetadata": { "identityData": { "identityProviderType": "Cognito", "issuer": "https://cognito-idp.us-east-1.amazonaws.com/", "sub": "c" } } } ], "humanLoopName": "test", "inputContent": { "document": { "documentType": "countryReport", "fields": [ { "id": "reportOwner", "type": "string", "validation": "", "value": "John Smith" }, { "id": "reportDate", "type": "date", "validation": "", "value": "2022-06-01T08:30" }, { "id": "locationList", "type": "table", "value": { "columns": [ { "id": "country", "type": "string" }, { "id": "capital", "type": "string" }, { "id": "population", "type": "number" } ], "rows": [ [ "UK", "London", 1234 ], [ "France", "Paris", 321 ] ] } } ] }, "document_types": [ { "displayName": "Email", "id": "email" }, { "displayName": "Invoice", "id": "invoice" }, { "displayName": "Other", "id": "other" } ], "input_s3_uri": "s3://my-input-bucket/file1.pdf" } }
期望输出格式
Input info: country, Original answer: UK, Human answer: extracted1_1: Italy Input info: capital, Original answer: London, Human answer: extracted1_2: Rome Input info: population, Original answer: 1234, Human answer: extracted1_3: 5555 Input info: country, Original answer: France, Human answer: extracted2_1: Czech Input info: capital, Original answer: Paris, Human answer: extracted2_2: Prague Input info: population, Original answer: 321, Human answer: extracted2_3: 3333
当前代码
s3_client = boto3.client('s3') response = s3_client.get_object(Bucket=f'{config["bucket"]}', Key=f'{config["file_name"]}') data = response['Body'].read() d = json.loads(data) column = d['inputContent']['document']['fields'][2]['value']['columns'] row = d['inputContent']['document']['fields'][2]['value']['rows'] answers = d['humanAnswers'][0]['answerContent'] str_row = str(row) iter_col = iter(column) iter_row = iter(str_row) combined = '' for a in answers.items(): nxt_col = next(iter_col) for list in row: for values in list: v = values combined += str(v + ", ") print(f'Input info: {nxt_col}, Original Answer: {str_row}, Human Answer: {a}')
解决方案
首先要理清数据对应逻辑:原始表格的2行数据,分别对应人工答案里的extracted1_*和extracted2_*组;每组里的_1/_2/_3依次对应表格的country/capital/population列。
修改后的代码如下:
import boto3 import json # 读取S3数据(假设config已定义) s3_client = boto3.client('s3') response = s3_client.get_object(Bucket=f'{config["bucket"]}', Key=f'{config["file_name"]}') data = response['Body'].read() d = json.loads(data) # 提取关键数据 columns = d['inputContent']['document']['fields'][2]['value']['columns'] original_rows = d['inputContent']['document']['fields'][2]['value']['rows'] human_answers = d['humanAnswers'][0]['answerContent'] # 整理列名列表 col_names = [col['id'] for col in columns] # 过滤并分组人工答案:只保留extracted开头的,按行号分组 extracted_answers = {} for key, value in human_answers.items(): if key.startswith('extracted'): row_num = key.split('_')[1] if row_num not in extracted_answers: extracted_answers[row_num] = [] extracted_answers[row_num].append((key, value)) # 按行号排序,确保顺序正确 sorted_rows = sorted(extracted_answers.keys()) # 遍历每一行数据,生成对比内容 for idx, row_num in enumerate(sorted_rows): original_row = original_rows[idx] human_row = extracted_answers[row_num] # 遍历每一列的对应数据 for col_idx, (col_name, original_val) in enumerate(zip(col_names, original_row)): human_key, human_val = human_row[col_idx] print(f'Input info: {col_name}, Original answer: {original_val}, Human answer: {human_key}: {human_val}') # 每行之间空一行 print()
代码说明
- 提取列名:从表格columns中取出每个列的id,得到
['country', 'capital', 'population'] - 分组人工答案:将
extracted1_*和extracted2_*分别归为两组,对应原始数据的两行 - 匹配对应关系:遍历每一行原始数据,同时对应一组人工答案,按列索引匹配列名、原始值和人工答案的键值对
- 生成输出:按照期望格式打印每条对比内容,行与行之间空一行
内容的提问来源于stack exchange,提问作者RunRabbit
相关产品推荐
相关产品推荐

