如何按相同值列分组读取CSV并统一SK对应的最大PR值
Solution for CSV to JSON with SK-specific Max PR Replacement
Got it, let's work through this problem together. Since you noted that each SK's maximum PR is always in its last row of data, we can leverage that to simplify our logic—no need to compute the max value from scratch for each SK. Here's a step-by-step approach using Python (a common tool for this kind of data manipulation):
Step 1: Capture Each SK's Maximum PR
First, we'll do a single pass through the CSV to record the final (and thus maximum) PR value for every SK. This works because the last row for each SK holds the highest PR.
Step 2: Rewrite Rows with Unified PR Values
Next, we'll read the CSV again, replace each row's PR with the pre-recorded maximum PR for its SK, then structure the data into the required JSON format.
Full Code Example
import csv import json # First pass: Collect the maximum PR for each SK (last row's PR) sk_max_pr = {} with open('input.csv', 'r', newline='', encoding='utf-8') as csv_input: csv_reader = csv.DictReader(csv_input) for row in csv_reader: # Overwrite the SK's PR each time we encounter it—last entry is the max sk_max_pr[row['SK']] = row['PR'] # Second pass: Replace PR values and build JSON-ready data output_data = [] with open('input.csv', 'r', newline='', encoding='utf-8') as csv_input: csv_reader = csv.DictReader(csv_input) for row in csv_reader: # Swap current PR with the SK's max PR row['PR'] = sk_max_pr[row['SK']] output_data.append(row) # Write the final JSON file with open('output.json', 'w', encoding='utf-8') as json_output: json.dump(output_data, json_output, indent=4, ensure_ascii=False)
What This Does
- For your sample CSV input:
SK,VR,ST,PR 1000,1000-Q1,10,187 1000,1000-Q2,20,925 1001,1001-Q1,10,112 1001,1001-Q2,30,120 - The resulting
output.jsonwill have every row for SK1000using PR925, and every row for SK1001using PR120:[ { "SK": "1000", "VR": "1000-Q1", "ST": "10", "PR": "925" }, { "SK": "1000", "VR": "1000-Q2", "ST": "20", "PR": "925" }, { "SK": "1001", "VR": "1001-Q1", "ST": "10", "PR": "120" }, { "SK": "1001", "VR": "1001-Q2", "ST": "30", "PR": "120" } ]
Notes
- If your CSV ever changes so that the max PR isn't in the last row, you can modify the first pass to calculate the actual maximum by converting
PRto an integer and comparing values:sk_max_pr[sk] = max(int(sk_max_pr.get(sk, 0)), int(row['PR'])) - Using
utf-8encoding ensures compatibility with special characters, andindent=4makes the JSON human-readable.
内容的提问来源于stack exchange,提问作者USM
相关产品推荐
相关产品推荐

