如何在Pandas DataFrame中根据权重区间匹配对应列的取值?
Solution for Interval-Based Data Lookup in Pandas DataFrame
Alright, let's tackle this problem head-on. The key challenge here is mapping your weight value to the correct interval column, where each column name acts as the upper bound of the previous interval. Let's break this down and build the solution step by step.
First, let's recap your scenario to ensure we're aligned:
- Your DataFrame uses
locationas its index. - Column names represent the upper bounds of intervals (e.g., '200', '1000', '2000').
- A weight value falls into the interval between the previous column's upper bound and the current column's upper bound (e.g., 500 sits in ]200;1000[, so we need to pull data from the '200' column).
Step 1: Set Up the Test DataFrame
First, let's recreate your sample DataFrame exactly as you described:
import pandas as pd # Initialize the test DataFrame data_frame_test = pd.DataFrame({ 'location': [1, 2, 'S'], '200': [342, 690, 103], '1000': [322, 120, 193], '2000': [249, 990, 403] }) # Set 'location' as the index data_frame_test = data_frame_test.set_index('location')
Step 2: Write the Lookup Function
This function will handle mapping the weight to the correct interval column, then fetch the corresponding value using the location index:
def get_weight_based_value(df, target_location, weight): # Convert column names from strings to integers and sort them sorted_col_bounds = sorted(int(col) for col in df.columns) # Handle edge case: weight is smaller than or equal to the smallest interval bound if weight <= sorted_col_bounds[0]: target_column = str(sorted_col_bounds[0]) else: # Find the first interval bound that's larger than the weight for bound in sorted_col_bounds: if bound > weight: # The target column is the previous bound in the sorted list target_column = str(sorted_col_bounds[sorted_col_bounds.index(bound) - 1]) break else: # Handle edge case: weight is larger than or equal to the largest interval bound target_column = str(sorted_col_bounds[-1]) # Retrieve the value using location index and target column return df.loc[target_location, target_column]
Step 3: Test the Function
Let's verify with your example case to make sure it works as expected:
# Test with location=1, weight=500 result = get_weight_based_value(data_frame_test, 1, 500) print(result) # Output: 342
Edge Case Handling
The function also covers these common scenarios:
- If
weight = 150(≤200), it returns the value from the '200' column for the given location. - If
weight = 1500(between 1000 and 2000), it returns the value from the '1000' column. - If
weight = 2500(≥2000), it returns the value from the '2000' column.
内容的提问来源于stack exchange,提问作者Murcielago
相关产品推荐
相关产品推荐

