如何实现列表元素到前一个零的间隔计数?(含Pandas问题代码)
Hey there! Let's get this sorted out. Your goal is to transform the list [7, 2, 0, 3, 4, 2, 5, 0, 3, 4] into [1, 2, 0, 1, 2, 3, 4, 0, 1, 2], where each value represents the interval to the previous zero (or the start of the list if there's no prior zero).
First, let's look at two clean, efficient ways to achieve this with Pandas, then we'll break down why your original code wasn't working.
Method 1: Group by Zero Occurrences & Cumulative Count
This approach uses grouping to track segments between zeros, then counts positions within each segment:
import pandas as pd # Initialize your DataFrame df = pd.DataFrame({'X': [7, 2, 0, 3, 4, 2, 5, 0, 3, 4]}) # Create a group ID: increments every time we hit a zero df['group_id'] = (df['X'] == 0).cumsum() # Calculate cumulative position within each group (start at 1) df['result'] = df.groupby('group_id').cumcount() + 1 # Set result to 0 where the original value is 0 df.loc[df['X'] == 0, 'result'] = 0 # Get the final list print(df['result'].tolist()) # Output: [1, 2, 0, 1, 2, 3, 4, 0, 1, 2]
How this works:
(df['X'] == 0).cumsum()generates a unique ID for each segment between zeros. For example, the first two non-zero values are in group 0, the values after the first zero are in group 1, etc.groupby('group_id').cumcount() +1counts the position of each element within its group (starting at 1 instead of 0).- Finally, we overwrite the results for zero values to 0 to match your requirement.
Method 2: Track Last Zero Position & Calculate Difference
This method uses forward filling to track the most recent zero index, then computes the interval as the difference between current index and last zero index:
import pandas as pd df = pd.DataFrame({'X': [7, 2, 0, 3, 4, 2, 5, 0, 3, 4]}) # Record the index of each zero; fill non-zero positions with NaN zero_indices = df.index.where(df['X'] == 0) # Forward fill to propagate the last zero index to subsequent positions # For values before the first zero, fill with -1 (so index - (-1) gives the correct interval) df['last_zero_idx'] = zero_indices.ffill().fillna(-1).astype(int) # Calculate interval: current index minus last zero index df['result'] = df.index - df['last_zero_idx'] # Set result to 0 where original value is 0 df.loc[df['X'] == 0, 'result'] = 0 print(df['result'].tolist()) # Output: [1, 2, 0, 1, 2, 3, 4, 0, 1, 2]
How this works:
df.index.where(df['X'] ==0)keeps only the indices whereXis 0, others become NaN.ffill()fills each NaN with the most recent non-NaN (i.e., last zero index).fillna(-1)handles the segment before the first zero.- The interval is simply the current index minus the last zero index (for the first segment, index 0 - (-1) =1, index1 - (-1)=2, which matches your requirement).
Why Your Original Code Failed
Let's go over some key issues in your code:
- Recursion is inefficient & error-prone: Recursive calls for each element aren't suited for Pandas, which is designed for vectorized operations. This also leads to messy index logic.
- Incorrect Series initialization:
target_series = pd.Series([])creates an empty Series, so assigning values to indices liketarget_series[actual_index]will cause errors (those indices don't exist yet). You should initialize it with the same length asdf_series, filled with default values. - Confusing index logic: The conditions like
i+df_series.size != actual_indexand handling negative indices don't align with the problem's requirements, leading to incorrect counts or missing values. - Unreliable assignment: Directly assigning to
target_series[actual_index]without initializing the Series properly means some positions might remain unassigned, resulting in NaN values.
Both methods above avoid these issues and leverage Pandas' strengths for clean, scalable code.
内容的提问来源于stack exchange,提问作者Bravo

