You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas筛选出现次数大于1的品种Value Counts统计结果?

How to Filter Wine Varieties with Occurrences > 1 Efficiently

Hey there! I get what you're trying to do—you want a clean, concise way to pull only those wine varieties that show up more than once in your dataset, along with their respective counts. Let's break down why your initial attempts hit errors, then jump straight to the simplest solution.

Why Your Previous Code Threw Errors

The core issue with lines like df.loc[df['variety'].value_counts()>1] is that df['variety'].value_counts() returns a Series where the index is the wine variety name, and the values are the occurrence counts. But df.loc expects row indices (not variety names) to filter the original DataFrame. These two index types don't match, hence the error. Same goes for your second attempt—you're trying to use variety names to index the original variety column, which uses row numbers as its index, so the alignment fails.

The Simplest, Most Efficient Solution

You don't need to loop back to the original DataFrame at all! The value_counts() result is already exactly what you need—you just have to filter that Series directly. Here's how:

# First get the full count of each wine variety
variety_counts = df['variety'].value_counts()

# Filter to keep only varieties with occurrence counts greater than 1
filtered_varieties = variety_counts[variety_counts > 1]

That's it! filtered_varieties will be a Series where each entry uses the variety name as its index, paired with its count—only for varieties that appear more than once. This is far more efficient than converting to a DataFrame and adding extra filter steps.

Bonus: Alternative with GroupBy (If You Need Full Row Data Later)

If you ever need to work with the actual rows of those filtered varieties down the line, you could use groupby + filter first, then get the counts—but this is a bit more verbose than the above method:

# Filter the original DataFrame to only include varieties with >1 occurrence
filtered_df = df.groupby('variety').filter(lambda x: len(x) > 1)

# Then get the counts from the filtered DataFrame
filtered_varieties = filtered_df['variety'].value_counts()

But for just grabbing the counts of frequent varieties, the first method is definitely the cleanest and fastest approach.

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:40:13