如何编写代码合并Pandas DataFrame中值总和≤5的坐标区间?
Solution to Merge Consecutive Intervals with Value Sum Limit
Here's a practical implementation that aligns with your requirements. The approach iterates through the intervals, accumulating values until adding the next would exceed 5 (or keeping any interval with a value over 5 as a standalone entry).
Code Implementation
import pandas as pd # Create the input DataFrame From = [10, 20, 30, 40, 50, 60, 70] to = [20, 30, 40, 50, 60, 70, 80] value = [2, 3, 5, 6, 1, 3, 1] df = pd.DataFrame({'from': From, 'to': to, 'value': value}) # Initialize variables to track merged intervals merged_intervals = [] current_start = None current_end = None current_sum = 0 # Iterate through each row to build merged intervals for _, row in df.iterrows(): current_val = row['value'] # Case 1: Current interval's value exceeds 5, keep it separate if current_val > 5: # Add any ongoing merged interval first if exists if current_sum > 0: merged_intervals.append({ 'from': current_start, 'to': current_end, 'value': current_sum }) # Reset tracking variables current_start = None current_end = None current_sum = 0 # Add the over-5 interval as a standalone entry merged_intervals.append({ 'from': row['from'], 'to': row['to'], 'value': current_val }) # Case 2: Current value is within limit, check if we can merge it else: if current_sum + current_val > 5: # Can't merge, finalize the current merged interval merged_intervals.append({ 'from': current_start, 'to': current_end, 'value': current_sum }) # Start a new merged interval with the current row current_start = row['from'] current_end = row['to'] current_sum = current_val else: # Start or continue the merged interval if current_sum == 0: current_start = row['from'] current_end = row['to'] else: current_end = row['to'] # Update end to current row's to current_sum += current_val # Add any remaining merged interval after loop ends if current_sum > 0: merged_intervals.append({ 'from': current_start, 'to': current_end, 'value': current_sum }) # Convert the merged list to a DataFrame result_df = pd.DataFrame(merged_intervals) print(result_df)
Output
from to value 0 10 30 5 1 30 40 5 2 40 50 6 3 50 80 5
Explanation
- Tracking Merged Intervals: We use
current_start,current_end, andcurrent_sumto keep tabs on the ongoing merged interval as we iterate. - Over-5 Values: Any interval with a value greater than 5 is immediately added as a standalone entry, and any in-progress merge is finalized first to avoid gaps.
- Merge Logic: For values ≤5, we check if adding the current value to the accumulated sum stays within the 5 limit. If not, we wrap up the current merge and start a new one. If yes, we extend the current merged interval to include the row's end coordinate.
- Final Cleanup: After the loop finishes, we add any remaining merged interval that wasn't finalized during iteration.
内容的提问来源于stack exchange,提问作者The Oracle
相关产品推荐
相关产品推荐

