如何按TaxRate分组相同父级Company的子级Item列表
解决方案:按TaxRate合并Company条目
我来帮你搞定这个按TaxRate合并Company条目的需求!先明确核心诉求:把原本按Month分散的同一Company+同一TaxRate的条目合并,将所有对应的Item整合到同一个Company节点下,最终每个(Company+TaxRate)组合只保留一个条目,Month字段可按需保留最后/第一个出现的版本。
实现思路
- 第一步:解析原始数据,拆分出每个独立的Company块(包含Company名称、TaxRate、Item列表、Month)
- 第二步:以
(Company名称, TaxRate)作为唯一分组键,将相同键的块中的Item列表合并 - 第三步:重新构建目标结构,每个分组生成一个完整的Company块
代码示例(Python)
假设原始数据是结构化字典
如果你的数据已经是结构化的(比如嵌套字典/对象),可以用以下代码处理:
from collections import defaultdict # 模拟原始结构化数据 business = { "Business": [ { "Company": "CompanyA", "TaxRate": 0.1, "Items": [ {"Name": "Item1", "Price": 100, "Total": 110}, {"Name": "Item2", "Price": 200, "Total": 220} ], "Month": "Jan" }, { "Company": "CompanyB", "TaxRate": 0.2, "Items": [ {"Name": "Item3", "Price": 150, "Total": 180}, {"Name": "Item4", "Price": 300, "Total": 360} ], "Month": "Jan" }, { "Company": "CompanyA", "TaxRate": 0.1, "Items": [ {"Name": "Item5", "Price": 120, "Total": 132}, {"Name": "Item6", "Price": 250, "Total": 275} ], "Month": "Feb" } ] } # 按(Company, TaxRate)分组合并Item grouped = defaultdict(lambda: {"Items": [], "Month": ""}) for company_block in business["Business"]: group_key = (company_block["Company"], company_block["TaxRate"]) grouped[group_key]["Items"].extend(company_block["Items"]) grouped[group_key]["Month"] = company_block["Month"] # 保留最后一个Month,可改为取第一个 # 重构目标结构 merged_business = { "Business": [ { "Company": comp_name, "TaxRate": tax_rate, "Items": item_list["Items"], "Month": item_list["Month"] } for (comp_name, tax_rate), item_list in grouped.items() ] } # 输出结果 import json print(json.dumps(merged_business, indent=2))
如果原始数据是纯文本格式
如果你的数据是用户提供的纯文本字符串,可以先解析再合并:
import re from collections import defaultdict # 模拟原始纯文本数据 raw_business_text = """Business -------- CompanyA TaxRate 0.1 Item1 Name Apple Price 100 Total 110 Item2 Name Banana Price 200 Total 220 Month Jan CompanyB TaxRate 0.2 Item3 Name Orange Price 150 Total 180 Item4 Name Grape Price 300 Total 360 Month Jan CompanyA TaxRate 0.1 Item5 Name Mango Price 120 Total 132 Item6 Name Pineapple Price 250 Total 275 Month Feb""" # 提取每个Company块的正则表达式 company_pattern = re.compile(r'(Company[A-Z]) TaxRate (\d+\.\d+) (.*?) Month (\w+)') company_blocks = company_pattern.findall(raw_business_text) # 分组合并Item grouped = defaultdict(lambda: {"Items": [], "Month": ""}) for comp_name, tax_rate, items_segment, month in company_blocks: # 提取单个Item信息 item_pattern = re.compile(r'Item(\d+) Name (\w+) Price (\d+) Total (\d+)') items = item_pattern.findall(items_segment) formatted_items = [ {"Name": f"Item{item_num}", "ProductName": name, "Price": int(price), "Total": int(total)} for item_num, name, price, total in items ] group_key = (comp_name, tax_rate) grouped[group_key]["Items"].extend(formatted_items) grouped[group_key]["Month"] = month # 重新构建目标文本 result_segments = ["Business --------"] for (comp_name, tax_rate), data in grouped.items(): segment_parts = [comp_name, f"TaxRate {tax_rate}"] for item in data["Items"]: segment_parts.extend([ item["Name"], f"Name {item['ProductName']}", f"Price {item['Price']}", f"Total {item['Total']}" ]) segment_parts.append(f"Month {data['Month']}") result_segments.append(' '.join(segment_parts)) final_result_text = ' '.join(result_segments) print(final_result_text)
运行后就能得到你想要的合并结构啦!
内容的提问来源于stack exchange,提问作者Triet Pham
相关产品推荐
相关产品推荐

