多CSV转无重复子节点JSON:D3格式转换问题求助
Fix Duplicate Child Nodes When Converting CSV to D3 Hierarchical JSON
Hey there! The core issue here is that your original code treats every CSV column (L1 to L6) as a mandatory hierarchical level—even when consecutive columns have identical values. This creates unnecessary nested duplicate nodes (like the repeated "young" levels). We need to adjust the logic to merge consecutive duplicate values and only attach the size to the final unique node in each hierarchy chain.
Modified Solution Code
import json import csv class Node(object): def __init__(self, name, size=None): self.name = name self.children = [] self.size = size def child(self, cname, size=None): child_found = [c for c in self.children if c.name == cname] if not child_found: _child = Node(cname, size) self.children.append(_child) else: _child = child_found[0] return _child def as_dict(self): res = {'name': self.name} if self.size is None: res['children'] = [c.as_dict() for c in self.children] else: res['size'] = self.size return res def remove_consecutive_duplicates(lst): """Remove consecutive duplicate values to avoid nested duplicate nodes""" if not lst: return [] unique_lst = [lst[0]] for item in lst[1:]: if item != unique_lst[-1]: unique_lst.append(item) return unique_lst root = Node('Segments') with open('C:\\Users\\G01172472\\Desktop\\Book3.csv', 'r') as f: reader = csv.reader(f) next(reader) # Skip the header row directly (more efficient than list conversion) for row in reader: # Extract hierarchy columns (first 6) and size value hierarchy_values = row[:6] size = row[6] # Clean up consecutive duplicates from the hierarchy chain cleaned_hierarchy = remove_consecutive_duplicates(hierarchy_values) # Traverse or create nodes using the cleaned hierarchy current_node = root for index, node_name in enumerate(cleaned_hierarchy): # Assign size only to the final node in the cleaned chain if index == len(cleaned_hierarchy) - 1: current_node = current_node.child(node_name, size) else: current_node = current_node.child(node_name) # Generate and print the desired JSON print(json.dumps(root.as_dict(), indent=4))
Key Changes Explained
- Consecutive Duplicate Removal: The
remove_consecutive_duplicatesfunction strips out repeated values in a row (e.g., turning["young", "young", "young", "young"]into["young"]). This eliminates the nested duplicate nodes entirely. - Header Handling: Replaced the list conversion of rows with
next(reader)to skip the header—this is more memory-efficient for large CSVs. - Targeted Size Assignment: We only attach the
sizevalue to the last node in the cleaned hierarchy chain, ensuring it's placed exactly where you need it instead of nested deep in duplicate levels.
When you run this code, it will output the exact hierarchical JSON structure you're expecting, with no duplicate child nodes.
内容的提问来源于stack exchange,提问作者Cyley Simon
相关产品推荐
相关产品推荐

