如何在Python中无需新建DataFrame获取各国Top10销售卖家?
Hey there! Let's fix this problem properly—your initial approaches either miss the mark entirely or are way more tedious than they need to be. Here's how to do this cleanly, without manually creating and concatenating multiple DataFrames.
Why Your First Approach Failed
The code df.loc[df['Country'].isin(['India','UK'])].sort_values(['Sales'], ascending=False)[0:20] doesn't work because it sorts all sellers from India and UK together, then takes the top 20. That means you could end up with 20 Indian sellers and zero UK sellers (or vice versa) if one country's top sellers outperform the other's—definitely not what you want when you need top 10 per country.
Why the Second Approach Isn't Ideal
While manually creating DataFrames for each country and concatenating them works, it's not scalable. If you add more countries later, you'll have to keep updating the code with new variables and concat calls. Even a loop is better, but we can do better than that.
The Clean, Scalable Solution
There are two great ways to achieve this—pick the one that fits your use case:
1. GroupBy + Apply (Most Readable)
This method groups the DataFrame by Country, sorts each group by Sales in descending order, and takes the top 10 entries from each group. It's straightforward and easy to follow:
top10_per_country = df.groupby('Country', group_keys=False).apply( lambda group: group.sort_values('Sales', ascending=False).head(10) )
group_keys=Falseprevents adding the group key (Country) as an extra index level, keeping your output clean.- This works for any number of countries—no need to update code if you add more regions later.
2. Rank + Filter (Better for Large Datasets)
If you're working with a huge DataFrame, this vectorized approach is more efficient. It calculates the sales rank per country, then filters to keep only entries where the rank is 10 or lower:
# Calculate sales rank within each country (1 = highest sales) df['sales_rank'] = df.groupby('Country')['Sales'].rank( ascending=False, method='first' # Handles ties by assigning unique ranks based on position ) # Filter to keep top 10 and drop the temporary rank column top10_per_country = df[df['sales_rank'] <= 10].drop('sales_rank', axis=1)
- The
method='first'ensures that if two sellers have the same sales, the one that appears first in the DataFrame gets a higher rank. You can switch tomethod='min'if you want all tied sellers to get the same rank (just note this might return more than 10 entries per country if there are ties at rank 10).
Testing with Your Sample Data
Using your provided sample DataFrame, both methods will return exactly 20 rows: 10 from India and 10 from UK, each sorted by highest sales first. No manual DataFrame juggling required!
内容的提问来源于stack exchange,提问作者Vivi

