如何用Pandas提取DataFrame列的唯一值、计数及索引键列
Solution to Extract Unique Values, Counts, and Index Keys from DataFrame Columns
Got it, let's walk through how to generate the desired DataFrame df2 from your example df1. Here's a step-by-step breakdown using pandas:
Step 1: Set Up the Example DataFrame
First, let's create the sample DataFrame you provided to work with:
import pandas as pd # Initialize df1 with your sample data data = { 'city': ['a', 'b', 'c', 'd', 'a', 'a'], 'score': [12, 45, 50, 65, 58, 70] } df1 = pd.DataFrame(data)
Step 2: Aggregate Unique Values, Counts, and Index Keys
We'll use groupby() to group by the city column, then aggregate to get both the count of each unique city and their corresponding original index positions:
# Group by city and aggregate count + index keys df2 = df1.groupby('city').agg( count=('city', 'size'), # Count occurrences of each city index_key=('city', lambda x: ', '.join(map(str, x.index))) # Collect original indices as a string ).reset_index().rename(columns={'city': 'unique city'}) # Optional: Add ellipsis for longer index lists (matches your example's "1 ..." format) df2['index key'] = df2['index_key'].apply(lambda x: x if len(x) <= 5 else f"{x.split(',')[0]} ...") df2 = df2.drop(columns='index_key') # Drop the intermediate column
Step 3: View the Result
Running the code above will produce df2 exactly as you need:
unique city count index key 0 a 3 0 ... 1 b 1 1 2 c 1 2 3 d 1 3
Quick Notes:
- If you want to keep the full list of indices without ellipsis, just remove the optional
apply()anddrop()lines. Theindex keycolumn will show all original indices separated by commas (e.g.,0, 4, 5for city 'a'). - The
sizemethod in the aggregation gives the count of rows per group, which is exactly what we need for thecountcolumn.
内容的提问来源于stack exchange,提问作者Gowtham Paramasivam
相关产品推荐
相关产品推荐

