Pandas条件计数实现:如何为车辆每日行程生成连续序号列
Hey there! Let's sort out this ride counting problem for you. First, let's break down why your current code isn't working, then I'll show you two reliable ways to get the exact result you want.
Why Your Current Code Fails
- Wrong parameter in
apply: When you usex[df["Is_first_ride"]], you're passing the entireIs_first_rideSeries to thecounterfunction instead of just the value from the current row. It should bex["Is_first_ride"]instead. - State isn't preserved across rows: The
counterfunction resetscounter = 1every time it's called (sinceapplyruns it row-by-row). Even if you fixed the parameter, the while loop would only run once per row, so you'd never get a running count that carries over between rows. - Ambiguous boolean check: Comparing a Series (from your incorrect parameter) to
1creates a boolean Series, which pandas can't interpret as a single True/False value—hence the "Series is Ambiguous" error.
Solution 1: Use Your Existing Is_first_ride Column
We can use the Is_first_ride flags to create grouping IDs, then count within each group:
# Create a unique group ID for each new day's first ride (each 1 in Is_first_ride starts a new group) df['group_id'] = df['Is_first_ride'].cumsum() # Generate the ride counter by counting positions within each group df['#_of_ride'] = df.groupby('group_id').cumcount() + 1 # Optional: Drop the temporary group_id column if you don't need it df = df.drop('group_id', axis=1)
Solution 2: Directly Group by Car and Date (Recommended)
Since you're counting rides per car per day, grouping directly by car and date is more straightforward and reliable (it doesn't depend on the Is_first_ride column being perfectly accurate):
# Group by car and date, then count the position of each row in its group (add 1 to start from 1) df['#_of_ride'] = df.groupby(['car', 'date']).cumcount() + 1
Test It With Your Sample Data
Here's a full example using your sample input to confirm it works:
import pandas as pd # Your sample data data = { 'car': ['Ford', 'Ford', 'Ford', 'Ford', 'Fiat', 'Fiat', 'Fiat', 'Fiat', 'Ford', 'Ford'], 'date': ['20.1.2021']*4 + ['20.1.2021']*4 + ['21.1.2021']*2, 'Is_first_ride': [1,0,0,0,1,0,0,0,1,0] } df = pd.DataFrame(data) # Apply Solution 2 (the simpler one) df['#_of_ride'] = df.groupby(['car', 'date']).cumcount() + 1 print(df)
This will output exactly the table you wanted:
car date Is_first_ride #_of_ride 0 Ford 20.1.2021 1 1 1 Ford 20.1.2021 0 2 2 Ford 20.1.2021 0 3 3 Ford 20.1.2021 0 4 4 Fiat 20.1.2021 1 1 5 Fiat 20.1.2021 0 2 6 Fiat 20.1.2021 0 3 7 Fiat 20.1.2021 0 4 8 Ford 21.1.2021 1 1 9 Ford 21.1.2021 0 2
内容的提问来源于stack exchange,提问作者alex_loxn
相关产品推荐
相关产品推荐

