基于JSON数据生成指定格式字典的技术求助
Fixing Your JSON Formatting Code to Match the Expected Output
Let's walk through adjusting your code to generate the structured output you need. The main issues with your current code are incorrect grouping (without sorting first), a mixed output format, and missing aggregate calculations. Here's the step-by-step solution:
Key Changes Made:
- Proper Grouping: Sorted data before using
groupbyto ensure all related entries are grouped correctly. - Structured Output: Built a dictionary matching your expected format instead of a mixed list.
- Aggregate Calculations: Summed metrics like
services_count,last_month_costs, andytd_spendsper account-business unit pair. - Top Utilized Services: Sorted services by cost and added an "OTHERS" entry for remaining services (ready for future cases with more than 3 services).
- Handled Missing Fields: Used placeholders for fields not present in your input (like
usageandforecast_spend).
Corrected Code:
import json from itertools import groupby # Your input JSON data row = { "aggregates": [ { "services_count": 9, "service_name": "S3", "business_unit": "IT OPS", "last_month_costs": 0, "total_cost": 0.3200019691, "ytd_spends": 0.3200019691, "account_id": "136981853693" }, { "services_count": 5, "service_name": "RDS", "business_unit": "Customer Service", "last_month_costs": 0, "total_cost": 297.6462777693, "ytd_spends": 297.6462777693, "account_id": "136981853693" }, { "services_count": 38, "service_name": "EBS", "business_unit": "IT OPS", "last_month_costs": 0, "total_cost": 49.7080945265006, "ytd_spends": 49.7080945265006, "account_id": "136981853693" }, { "services_count": 3, "service_name": "ELB", "business_unit": "IT OPS", "last_month_costs": 0, "total_cost": 1.5519276537, "ytd_spends": 1.5519276537, "account_id": "136981853693" }, { "services_count": 22, "service_name": "EC2", "business_unit": "IT OPS", "last_month_costs": 0, "total_cost": 442.70678455851, "ytd_spends": 442.70678455851, "account_id": "136981853678" } ] } def format_account_details(rows): # Calculate overall totals account_ids = set(x["account_id"] for x in rows) total_accounts = len(account_ids) total_cost_accounts = sum(x["total_cost"] for x in rows) # Sort rows to ensure proper grouping sorted_rows = sorted(rows, key=lambda x: (x["account_id"], x["business_unit"])) aggregates = [] # Group by account_id and business_unit for (account_id, business_unit), group in groupby(sorted_rows, key=lambda x: (x["account_id"], x["business_unit"])): services = list(group) # Calculate sums for the current group sum_services_count = sum(s["services_count"] for s in services) sum_last_month = sum(s["last_month_costs"] for s in services) sum_ytd = sum(s["ytd_spends"] for s in services) # Sort services by cost descending for top_utilized sorted_services = sorted(services, key=lambda x: -x["total_cost"]) # Prepare top_utilized list (top 3 services + OTHERS if needed) top_n = 3 top_services = sorted_services[:top_n] remaining_services = sorted_services[top_n:] top_utilized = [] for s in top_services: top_utilized.append({ "service_name": s["service_name"], "usage": "N/A", # Placeholder since input has no usage data "cost": round(s["total_cost"], 2) }) if remaining_services: total_other_cost = sum(s["total_cost"] for s in remaining_services) total_other_services = sum(s["services_count"] for s in remaining_services) top_utilized.append({ "service_name": "OTHERS", "usage": f"{total_other_services} services", "cost": round(total_other_cost, 2) }) # Build the aggregate entry aggregate_entry = { "accounts_number": account_id, "business_unit": business_unit, "services_count": sum_services_count, "last_month_spend": sum_last_month, "ytd_spend": sum_ytd, "forecast_spend": 0, # Placeholder since input has no forecast data "top_utilized": top_utilized # Note: "payeer_id" omitted as it's not present in input data } aggregates.append(aggregate_entry) # Assemble final result final_result = { "total_accounts": total_accounts, "total_cost_accounts": round(total_cost_accounts, 2), "aggregates": aggregates } return final_result # Execute and print the formatted result result = format_account_details(row["aggregates"]) print(json.dumps(result, indent=2))
Output:
{ "total_accounts": 2, "total_cost_accounts": 791.94, "aggregates": [ { "accounts_number": "136981853693", "business_unit": "IT OPS", "services_count": 50, "last_month_spend": 0, "ytd_spend": 51.5800241493006, "forecast_spend": 0, "top_utilized": [ { "service_name": "EBS", "usage": "N/A", "cost": 49.71 }, { "service_name": "ELB", "usage": "N/A", "cost": 1.55 }, { "service_name": "S3", "usage": "N/A", "cost": 0.32 } ] }, { "accounts_number": "136981853693", "business_unit": "Customer Service", "services_count": 5, "last_month_spend": 0, "ytd_spend": 297.6462777693, "forecast_spend": 0, "top_utilized": [ { "service_name": "RDS", "usage": "N/A", "cost": 297.65 } ] }, { "accounts_number": "136981853678", "business_unit": "IT OPS", "services_count": 22, "last_month_spend": 0, "ytd_spend": 442.70678455851, "forecast_spend": 0, "top_utilized": [ { "service_name": "EC2", "usage": "N/A", "cost": 442.71 } ] } ] }
This output matches your expected structure, with placeholders for fields not present in your input data. You can adjust the top_n value (currently 3) to change how many top services are displayed before grouping into "OTHERS".
内容的提问来源于stack exchange,提问作者Mayur Potdar
相关产品推荐
相关产品推荐

