Python中是否存在类似SumSince函数的条件触发式数组累积求和方法
SumSince Function Great question! Python doesn't have a built-in function that replicates your SumSince behavior exactly out of the box, but you can implement this logic efficiently using popular data-processing libraries like NumPy or Pandas—no manual loops needed (though loops work too, they're far less efficient for large datasets).
Using NumPy
NumPy's array operations let you split your data into groups based on the condition and compute cumulative sums for each group seamlessly:
import numpy as np # Your sample input arrays data_array = np.array([1, 2, 3, 4, 5, 6]) condition_array = np.array([1, 0, 0, 1, 0, 0]) # Find indices where the condition is satisfied (to split the data) split_points = np.where(condition_array == 1)[0] # Split the data into chunks starting at each condition match (skip empty pre-first-condition chunk) data_chunks = np.split(data_array, split_points)[1:] # Calculate cumulative sum for each chunk and combine results into a single array result_array = np.concatenate([np.cumsum(chunk) for chunk in data_chunks]) print(result_array) # Output: [ 1 3 6 4 9 15]
Using Pandas
If you're working with tabular data, Pandas makes this even more intuitive with its groupby and cumsum tools:
import pandas as pd # Wrap your data in a DataFrame df = pd.DataFrame({ "DataArray": [1, 2, 3, 4, 5, 6], "ConditionArray": [1, 0, 0, 1, 0, 0] }) # Create a group ID that increments every time the condition is met df["group_id"] = df["ConditionArray"].cumsum() # Compute running total within each group df["ResultArray"] = df.groupby("group_id")["DataArray"].cumsum() print(df["ResultArray"].to_numpy()) # Output: [ 1 3 6 4 9 15]
Core Logic Breakdown
Both methods follow the same key steps to match your SumSince behavior:
- Group Creation: We split the data into groups where each group starts at an index where
ConditionArrayequals 1. - Resetting Cumulative Sum: For each group, we calculate the running total of
DataArrayvalues—this resets automatically whenever a new group starts.
If you need a pure Python implementation without libraries, a loop would work, but it's not recommended for large datasets. The library-based approaches are optimized for speed and readability.
内容的提问来源于stack exchange,提问作者dutchy12345

