如何将年龄列中的区间值替换为对应中间值并去除特殊符号?
Convert Interval Strings in Age Column to Midpoints
Got it, let's solve this problem cleanly! You need to turn those bracket-enclosed age intervals into their midpoints while stripping out the [, -, and ) characters. Here's a straightforward way to do this with pandas, which is perfect for batch processing column data:
Step 1: Set up your data (example)
First, let's assume you're working with a pandas DataFrame. Here's sample data matching your interval format:
import pandas as pd # Sample DataFrame with your Age interval format df = pd.DataFrame({ 'Age': ['[0-10)', '[10-20)', '[20-30)', '[30-40)', '[40-50)', '[50-60)', '[60-70)', '[70-80)'] })
Step 2: Extract interval bounds and calculate midpoints
We'll use regular expressions to pull out the lower and upper bounds of each interval, then compute their average (the midpoint):
# Extract lower and upper numbers from each interval string df[['lower_bound', 'upper_bound']] = df['Age'].str.extract(r'(\d+)-(\d+)', expand=True) # Convert the extracted bounds from strings to integers df[['lower_bound', 'upper_bound']] = df[['lower_bound', 'upper_bound']].astype(int) # Calculate the midpoint for each interval df['Age'] = (df['lower_bound'] + df['upper_bound']) / 2 # Optional: Drop the temporary bound columns if you don't need them df = df.drop(['lower_bound', 'upper_bound'], axis=1)
What this does:
- The regex
r'(\d+)-(\d+)'targets the numeric values on either side of the-in each interval, capturing them as separate columns. - Converting these captures to integers lets us do arithmetic to find the midpoint (average of lower and upper bounds).
- We overwrite the original
Agecolumn with the midpoints, so you end up with clean numeric values instead of interval strings.
After running this, your Age column will have values like 5.0, 15.0, 25.0, etc.—exactly the midpoints you want!
内容的提问来源于stack exchange,提问作者Vish10
相关产品推荐
相关产品推荐

