如何使用Pandas Python默认函数计算布尔序列中False前后True值的求和?
Let's break down your two Pandas questions with practical, code-first solutions—no fancy hacks, just native functions you already have access to.
Q1: Summing a continuous True/False boolean sequence in Pandas
First, a quick reminder: in Pandas, True is implicitly treated as 1 and False as 0. So if you just need the total number of True values across the entire sequence, it's dead simple:
import pandas as pd bool_series = pd.Series([True, True, False, True, False, True, True]) total_true_count = bool_series.sum() print(total_true_count) # Output: 5
If you're actually looking to get the length of each consecutive block of True values (e.g., "how many Trues in a row before a False hits?"), use a grouping trick with a cumulative sum on the inverse of your boolean series:
# Create a group ID: every False increments the ID, grouping consecutive Trues together group_ids = (~bool_series).cumsum() # Group by the ID and sum (this gives the count of Trues per continuous block) continuous_true_lengths = bool_series.groupby(group_ids).sum() # Filter out groups that correspond to False blocks (their sum will be 0) continuous_true_lengths = continuous_true_lengths[continuous_true_lengths > 0] print(continuous_true_lengths) # Output: # 0 2 # 1 1 # 3 2 # dtype: int64
Q2: Calculating sums of True-associated values separated by False (e.g., 3+1+5=9, 2+6=8)
From your example, it looks like you have numeric values paired with a boolean flag, where False acts as a separator between chunks of True values. You want to sum the numbers in each of these separated chunks. Here's how to do this with native Pandas tools:
First, let's set up sample data that matches your example:
df = pd.DataFrame({ 'values': [3, 1, 5, 2, 6], 'is_true': [True, True, True, False, True, True] # False splits the two True chunks })
Now follow these steps:
- Generate a group ID using cumulative sum of the inverse boolean column. Each
Falsewill increment the ID, so all subsequentTruevalues get grouped together. - Filter out rows where
is_trueisFalse(we don't need those in our sums). - Group by the ID and sum the
valuescolumn.
Here's the code:
# Create group identifiers df['group_id'] = (~df['is_true']).cumsum() # Calculate sums for each True-only group chunk_sums = df[df['is_true']].groupby('group_id')['values'].sum() print(chunk_sums) # Output: # group_id # 0 9 # 1 8 # Name: values, dtype: int64
This gives exactly the sums you mentioned: 3+1+5=9 for the first chunk, 2+6=8 for the second. If your use case is a bit different (like summing values immediately before AND after a single False), feel free to clarify—but this approach fits your example perfectly.
内容的提问来源于stack exchange,提问作者id101112

