如何更新字典:更新已有值并追加新值(含DataFrame场景)
Great question! Let’s break this down into two clear parts: first the general dictionary update logic, then the specific pandas DataFrame scenario you’re working through.
General Dictionary Update Logic
At its core, updating a dictionary to sum existing values and add new entries just requires checking each key from your new data: if the key exists, add the new value to the existing one; if not, create a new key with the new value. Here are two straightforward ways to do this:
Method 1: Explicit Loop with dict.get()
This is easy to follow, even if you’re new to Python:
my_dict = {'key1': 21, 'key2': 15} new_data = {'key1': 18, 'key3': 7} for key, value in new_data.items(): # Use get() to safely retrieve existing value (or 0 if key doesn't exist) my_dict[key] = my_dict.get(key, 0) + value print(my_dict) # Output: {'key1': 39, 'key2': 15, 'key3': 7}
Method 2: Using collections.Counter
If you’re working with numeric counts (like your use case), Counter from the collections module is purpose-built for this kind of aggregation:
from collections import Counter my_counter = Counter({'key1': 21, 'key2': 15}) new_data = Counter({'key1': 18, 'key3': 7}) # The update() method automatically sums existing keys and adds new ones my_counter.update(new_data) print(dict(my_counter)) # Output: {'key1': 39, 'key2': 15, 'key3': 7}
Your Specific Pandas DataFrame Scenario
Now let’s apply this to your task: aggregating value_counts() results from multiple DataFrames into one dictionary, with summed counts for matching keys and new keys for unique entries.
Solution 1: Using Counter (Simplest Approach)
Since value_counts() returns a pandas Series that converts cleanly to a Counter, this is the most concise method:
import pandas as pd from collections import Counter # Initialize an empty Counter to hold all aggregated counts total_counts = Counter() # Replace this with your actual list of N DataFrames dataframes = [ pd.DataFrame({'col': ['key1', 'key1', 'key2']}), pd.DataFrame({'col': ['key1', 'key3', 'key3', 'key3']}), pd.DataFrame({'col': ['key2', 'key2', 'key4']}) ] # Loop through each DataFrame and update the total counts for df in dataframes: df_counts = Counter(df['col'].value_counts().to_dict()) total_counts.update(df_counts) # Convert back to a regular dictionary if needed final_dict = dict(total_counts) print(final_dict) # Output: {'key1': 3, 'key2': 3, 'key3': 3, 'key4': 1}
Solution 2: Using a Regular Dictionary (No Extra Imports)
If you prefer not to use collections, you can use the dict.get() method from earlier:
import pandas as pd # Initialize empty dictionary for aggregated counts total_counts = {} dataframes = [ pd.DataFrame({'col': ['key1', 'key1', 'key2']}), pd.DataFrame({'col': ['key1', 'key3', 'key3', 'key3']}), pd.DataFrame({'col': ['key2', 'key2', 'key4']}) ] for df in dataframes: # Get value counts as a dictionary df_counts = df['col'].value_counts().to_dict() # Update the total dictionary for key, count in df_counts.items(): total_counts[key] = total_counts.get(key, 0) + count print(total_counts) # Same output: {'key1': 3, 'key2': 3, 'key3': 3, 'key4': 1}
Both methods work perfectly for your scenario—pick whichever you find more readable!
内容的提问来源于stack exchange,提问作者Vash

