You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用groupby移除price_per_sqft离群值后,如何清洗位置名称?

问题:如何清理location字段中的冗余内容?

我正在开展一项数据清洗项目,需要移除price_per_sqft字段的离群值。为此我使用groupby函数按location分组,通过均值与标准差筛选出无离群值的子数据框,再用concat合并得到结果数据框。但输出结果的location字段中包含“1st Phase”“1st Block”这类冗余内容,如何获取干净的位置名称?

用户代码:

def remove_pps_outliers(df):
    df_out = pd.DataFrame()
    for key, subdf in df.groupby('location'):
        m = np.mean(subdf.price_per_sqft)
        st = np.std(subdf.price_per_sqft)
        reduced_df = subdf[(subdf.price_per_sqft>(m-st)) & (subdf.price_per_sqft<=(m+st))]
        df_out = pd.concat([df_out,reduced_df],ignore_index=True)
    return df_out
df6 = remove_pps_outliers(df5)
df6.head()

解决方案

方法1:正则表达式批量清理冗余前缀/后缀

针对“1st Phase”“2nd Block”这类带数字和固定关键词的冗余内容,用正则匹配后直接替换为空:

# 匹配数字+序数词(可选)+空格+Phase/Block/Layout等关键词,替换为空
df6['clean_location'] = df6['location'].str.replace(r'\d+(st|nd|rd|th)?\s+(Phase|Block|Layout|Sector)', '', regex=True).str.strip()

如果冗余内容在字段末尾,只需调整正则表达式的匹配逻辑即可。

方法2:按关键词分割提取核心地名

如果冗余部分都是固定的关键词(比如Phase、Block),可以用分割的方式提取核心部分:

# 按冗余关键词分割,取分割后的最后一段(假设核心地名在后面)
df6['clean_location'] = df6['location'].str.split(r'\s+(Phase|Block|Layout)\s+', expand=True)[1].fillna(df6['location'])

如果核心地名在字段开头,把索引从[1]改成[0]即可。

方法3:建立映射表精准替换

如果清楚每个冗余地名对应的干净名称,可以先整理映射字典,再批量替换:

location_mapping = {
    '1st Phase JP Nagar': 'JP Nagar',
    '1st Block Koramangala': 'Koramangala',
    # 补充其他需要替换的对应关系
}
df6['clean_location'] = df6['location'].map(location_mapping).fillna(df6['location'])

方法4:整合到离群值处理函数中同步完成

可以把清理逻辑直接加入现有函数,避免后续重复处理数据:

def remove_pps_outliers_and_clean_location(df):
    df_out = pd.DataFrame()
    
    # 定义location清理函数
    def clean_loc(loc):
        return loc.replace(r'\d+(st|nd|rd|th)?\s+(Phase|Block)', '', regex=True).strip()
    
    for key, subdf in df.groupby('location'):
        m = np.mean(subdf.price_per_sqft)
        st = np.std(subdf.price_per_sqft)
        reduced_df = subdf[(subdf.price_per_sqft>(m-st)) & (subdf.price_per_sqft<=(m+st))]
        # 为当前分组添加清理后的location字段
        reduced_df['clean_location'] = clean_loc(key)
        df_out = pd.concat([df_out, reduced_df], ignore_index=True)
    return df_out

df6 = remove_pps_outliers_and_clean_location(df5)

内容的提问来源于stack exchange,提问作者Ruchit Sheta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 22:18:29