如何根据字典数据替换Pandas DataFrame中的count列值?
Solution for Replacing DataFrame Count Values Based on Dictionary Tuples
Here's a straightforward way to achieve exactly what you need. We'll cover two methods—one optimized for readability (great for small datasets) and another for performance (ideal for larger datasets):
Method 1: Using apply and dict.get() (Readable for Small Data)
This approach checks each row's (A,B) pair against the dictionary, replacing the count value only when the pair exists as a key:
import pandas as pd # Your initial data s_dict = {('A1','B1'):100, ('A3','B3'):300} df = pd.DataFrame(data={'A': ['A1', 'A2'], 'B': ['B1', 'B2'], 'C': ['C1', 'C2'], 'count':[1,2]}) # Update the count column df['count'] = df.apply(lambda row: s_dict.get((row['A'], row['B']), row['count']), axis=1) print(df)
Output:
A B C count 0 A1 B1 C1 100 1 A2 B2 C2 2
How this works:
- The
lambdafunction runs on every row of the DataFrame s_dict.get((row['A'], row['B']), row['count'])looks up the (A,B) tuple in the dictionary. If the tuple exists as a key, we use the corresponding dictionary value; if not, we keep the original count value.
Method 2: Using map with Tuples (Efficient for Large Data)
For bigger datasets, this vectorized approach is faster because it avoids slow row-wise operations:
import pandas as pd # Your initial data s_dict = {('A1','B1'):100, ('A3','B3'):300} df = pd.DataFrame(data={'A': ['A1', 'A2'], 'B': ['B1', 'B2'], 'C': ['C1', 'C2'], 'count':[1,2]}) # Create a series of (A,B) tuples matching the dictionary's key format ab_tuples = df[['A', 'B']].apply(tuple, axis=1) # Map tuples to the dictionary, fill missing values with original count df['count'] = ab_tuples.map(s_dict).fillna(df['count']).astype(int) print(df)
Output is identical to the first method.
How this works:
ab_tuplesconverts the A and B columns into a Series of tuples that match the dictionary's key structuremap(s_dict)replaces tuples with their corresponding dictionary values (returns NaN for tuples that aren't keys)fillna(df['count'])restores the original count for missing keys, andastype(int)ensures we keep integer values (since NaN temporarily converts the column to float)
Quick Notes:
- If you don't want to modify the original DataFrame, create a copy first:
df_copy = df.copy()then make changes todf_copy - Pick Method 1 for small datasets (easier to read and debug) and Method 2 for large datasets (faster execution)
内容的提问来源于stack exchange,提问作者Shweta
相关产品推荐
相关产品推荐

