如何用Pandas基于列条件将当前值与前值拼接?
Hey there! Since you're new to Python and Pandas, let's walk through exactly how to replicate that Excel task you described. It's a common cumulative concatenation problem with a reset condition, and Pandas has efficient ways to handle this without messy row-by-row loops.
Step 1: Set up your sample data
First, let's create the DataFrame from your example to test our solution:
import pandas as pd # Your sample data data = { 'A': [False, True, True, False], 'B': ['bird', 'fish', 'Tiger', 'Elephant'] } df = pd.DataFrame(data)
Step 2: Create group identifiers
The key here is to group rows together where we need to concatenate values. Every time column A is False, we start a new group (since that's where we reset the concatenation). We can generate these groups using a cumulative sum:
# Generate group IDs: increments every time A is False df['group'] = df['A'].eq(False).cumsum()
Step 3: Cumulative concatenation within each group
Now we'll use groupby to process each group separately. For each group, we'll build the concatenated string incrementally using itertools.accumulate (this is more efficient than looping for larger datasets):
from itertools import accumulate # Define how to concatenate the previous value with the current B value def concat_with_comma(prev, current): return f"{prev},{current}" # Apply cumulative concatenation to each group df['C'] = df.groupby('group')['B'].transform( lambda group: list(accumulate(group, concat_with_comma)) )
Step 4: Check the result
If you print the DataFrame now, you'll see it matches your expected output perfectly:
print(df)
Output:
A B group C 0 False bird 1 bird 1 True fish 1 bird,fish 2 True Tiger 1 bird,fish,Tiger 3 False Elephant 2 Elephant
How this works
- Group IDs:
df['A'].eq(False).cumsum()creates a unique number for each "block" of rows starting with aFalsein columnA. All subsequentTruerows stay in the same group until the nextFalse. - Cumulative concatenation:
accumulateiterates through each group'sBvalues, taking the previous concatenated string and appending the currentBvalue with a comma. This mimics exactly what you'd do manually in Excel when dragging a formula down.
Optional: Remove the temporary group column
If you don't need the group column in your final output, you can drop it with:
df = df.drop('group', axis=1)
内容的提问来源于stack exchange,提问作者TVD

