You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化含重复try-except结构的DataFrame数据处理代码?

代码优化方案

基础优化版(保留循环,消除冗余)

把重复的异常处理逻辑封装成通用函数,减少代码冗余同时优化查询效率:

import numpy as np
import pandas as pd

# 封装坐标转换的安全处理函数
def safe_float_conversion(series):
    try:
        return float(series.unique())
    except:
        return np.nan

# 封装计数的安全处理函数
def safe_count(df, city):
    try:
        return df[df['City'] == city].shape[0]
    except:
        return 0.0001

# 获取去重城市列表(保留原顺序)
cities = nlp_inst.raw_data_frames['SKY_modified']['City'].unique().tolist()
mapping_data = []

for city in cities:
    # 提前取出当前城市的正向数据,避免重复过滤
    pos_city_subset = pos_df[pos_df['City'] == city]
    
    lat = safe_float_conversion(pos_city_subset['Latitude'])
    long = safe_float_conversion(pos_city_subset['Longitude'])
    
    city_count_pos = safe_count(pos_df, city)
    city_count_neg = safe_count(neg_df, city)
    
    ratio = city_count_pos / (city_count_pos + city_count_neg)
    mapping_data.append([lat, long, ratio])
    print(mapping_data[-1])

进阶优化版(向量操作,提升效率)

如果数据量较大,用pandas的向量化操作替代循环,运行效率会显著提升:

import numpy as np
import pandas as pd

# 获取去重城市列表
cities = nlp_inst.raw_data_frames['SKY_modified']['City'].unique().tolist()

# 批量统计正负样本的城市数量,不存在的城市填充默认值0.0001
pos_counts = pos_df['City'].value_counts().reindex(cities, fill_value=0.0001)
neg_counts = neg_df['City'].value_counts().reindex(cities, fill_value=0.0001)

# 批量获取城市坐标,空值转nan
city_coords = pos_df.groupby('City').agg(
    Latitude=('Latitude', 'first'),
    Longitude=('Longitude', 'first')
).reindex(cities)
city_coords['Latitude'] = city_coords['Latitude'].apply(lambda x: float(x) if pd.notna(x) else np.nan)
city_coords['Longitude'] = city_coords['Longitude'].apply(lambda x: float(x) if pd.notna(x) else np.nan)

# 计算比例并转换为目标格式
city_coords['ratio'] = pos_counts / (pos_counts + neg_counts)
mapping_data = city_coords[['Latitude', 'Longitude', 'ratio']].values.tolist()

for item in mapping_data:
    print(item)

优化说明

  • 消除冗余代码:把重复的try-except逻辑封装成函数,后续修改异常处理逻辑只需改动函数,不用逐个修改循环内的代码。
  • 提升查询效率:基础版中提前取出当前城市的子集,避免多次重复过滤DataFrame;进阶版直接用pandas向量操作,彻底替代Python循环,数据量越大优势越明显。
  • 更合理的城市列表生成:原代码用list(set(...))会打乱城市顺序,改用unique().tolist()既能去重又保留原始顺序。
  • 可靠的计数方式:用shape[0](行数)替代count()[0],避免因列中存在空值导致计数不准的问题。

内容的提问来源于stack exchange,提问作者jack gell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 18:30:18