Python中基于pandas DataFrame两列计算生成新列并处理NaN的方法
自定义计算财务指标并追加到DataFrame列的解决方案
问题根因
你当前代码无法正常生成EV/Ratio列的核心问题如下:
- 当
e = float(b)/float(a)计算触发异常时,你仅在except块中写了None,没有给e变量赋值,此时e未定义会触发程序报错,即使报错被忽略也无法得到合法空值填入col_e,最终会导致各列长度不匹配,无法正常生成DataFrame。 - 空的
except:语句会捕获所有类型异常,不利于后续排查其他潜在问题。 - 未处理ebitda为0的除零场景,以及b或a为None时的类型转换异常。
修改后完整代码
import time import yfinance as yf import concurrent.futures import pandas as pd # 此处替换为你自己的tickers列表 tickers = ["A", "AA", "AAC", "AACG", "AACIU", "AADI", "AAIC"] start = time.time() col_a = [] col_b = [] col_c = [] col_d = [] col_e = [] print('Loading Data... Please wait for results') def do_something(ticker): print('---', ticker, '---') all_info = yf.Ticker(ticker).info # 先给所有变量设置默认值,避免未定义错误 a = b = c = d = e = None try: a = all_info.get('ebitda') b = all_info.get('enterpriseValue') c = all_info.get('trailingPE') d = all_info.get('sector') # 校验有效值再计算,避免无意义的异常 if b is not None and a is not None and a != 0: e = float(b) / float(a) # 只捕获明确需要处理的异常类型,避免误吞其他错误 except (TypeError, ZeroDivisionError, ValueError): # 异常时e保持默认None即可 pass col_a.append(a) col_b.append(b) col_c.append(c) col_d.append(d) col_e.append(e) return with concurrent.futures.ThreadPoolExecutor() as executer: executer.map(do_something, tickers) # Dataframe Set Up pd.set_option("display.max_rows", None) df = pd.DataFrame({ 'Ticker': tickers, 'Ebitda': col_a, 'EnterpriseValue' :col_b, 'PE Ratio': col_c, 'Sector': col_d, 'EV/Ratio': col_e }) # 输出全量数据 print(df) print("===== 去空值后结果 =====") print(df.dropna()) print('It took', time.time()-start, 'seconds.')
关键修改说明
- 进入try块前先给所有变量设置默认None值,不管计算是否成功,各列追加的元素数量都能保持一致,不会出现DataFrame构造失败的问题。
- 计算EV/Ratio前增加了有效性校验,只有enterpriseValue和ebitda都有值且ebitda不为0时才执行计算,减少不必要的异常触发。
- 限定了异常捕获的类型,避免把接口请求失败、权限错误等其他问题也一并忽略,方便后续排查问题。
- 该逻辑可以直接扩展到其他自定义指标的计算,只需要在函数中新增对应的变量和计算逻辑,再把变量追加到对应的列表,最后添加到DataFrame的构造参数中即可。
内容的提问来源于stack exchange,提问作者Trevor Seibert
相关产品推荐
相关产品推荐

