含NaN值的Pandas DataFrame列双Y轴绘图异常求助
解决Matplotlib TwinX双Y轴含NaN数据的绘图问题
我来帮你搞定这个双Y轴绘图的问题!你遇到的核心问题是:第二列数据中大量的NaN导致matplotlib无法画出连续线条,且默认没有标记离散的有效数据点,所以看起来像是没显示。下面给你具体的解决方案和修改后的代码:
问题原因分析
matplotlib的plot()函数默认会绘制连续线条,但当数据中存在大量连续NaN时,它会中断线条;而你的第二列只有每10年一个有效数据,其余全是NaN,既没有连续数值支撑线条,又没设置数据点标记,所以图表上看不到第二列的内容。另外,你的第一列数据是字符串类型,建议先转为数值型,避免潜在的绘图异常。
解决方案
这里有两种实用的处理方式,你可以根据需求选择:
方式1:提取非NaN数据单独绘图(推荐)
直接筛选出第二列中有效的非NaN数据,绘制带标记的点(或线条),这样能清晰展示离散的时间点数据:
import pandas as pd import numpy as np import matplotlib.pyplot as plt # 原始数据 list1 = ['1297606', '1300760', '1303980', '1268987', '1333521', '1328570', '1328112', '1353671', '1371285', '1396658', '1429247', '1388937', '1359145', '1330414', '1267415', '1210883', '1221585', '1186039', '884273', '861789', '857475', '853485', '854122', '848163', '839226', '820151', '852385', '827609', '825564', '789217', '765651'] list1a = [1980, 1981, 1982, 1983, 1984, 1985, 1986, 1987, 1988, 1989, 1990, 1991, 1992, 1993, 1994, 1995, 1996, 1997, 1998, 1999, 2000, 2001, 2002, 2003, 2004, 2005, 2006, 2007, 2008, 2009, 2010] list3b = [121800016.0, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, 145279588.0, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, 160515434.5, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, 168140487.0] # 创建DataFrame并修正数据类型 d = {'Year': list1a,'Abortions per Year': list1, 'Affiliation with Religious Institutions': list3b} newdf = pd.DataFrame(data=d) newdf.set_index('Year',inplace=True) # 将第一列转为整数类型 newdf['Abortions per Year'] = newdf['Abortions per Year'].astype(int) # 绘图部分 fig, ax1 = plt.subplots(figsize=(20,5)) # 绘制第一列数据 ax1.plot(newdf['Abortions per Year'], color='blue', label='Abortions per Year') ax1.set_ylabel('Abortions Count', color='blue') ax1.tick_params(axis='y', labelcolor='blue') # 创建双Y轴并绘制第二列非NaN数据 ax1b = ax1.twinx() # 提取非NaN的行 non_nan_affil = newdf['Affiliation with Religious Institutions'].dropna() # 绘制带标记的点,可选添加线条连接 ax1b.plot(non_nan_affil.index, non_nan_affil.values, marker='o', linestyle='-', color='orange', markersize=8, label='Religious Affiliation') # 设置第二Y轴范围和标签 y2min = non_nan_affil.min() y2max = non_nan_affil.max() ax1b.set_ylim(y2min*0.8, y2max*1.2) ax1b.set_ylabel('Religious Institution Affiliation', color='orange') ax1b.tick_params(axis='y', labelcolor='orange') # 添加图例 fig.legend(loc='upper right', bbox_to_anchor=(1.1, 1)) plt.title('Abortions vs Religious Affiliation Over Time') plt.show()
方式2:直接修改绘图参数显示离散点
如果你不想筛选数据,也可以直接在plot()中设置标记样式,让matplotlib显示出离散的有效数据点:
# 其他代码和上面一致,仅修改ax1b的plot部分 ax1b.plot(newdf['Affiliation with Religious Institutions'], marker='s', linestyle='--', color='red', markersize=8, label='Religious Affiliation')
这种方式会保留原数据的索引,但因为中间都是NaN,线条会断开,不过标记点会清晰显示出来。
额外提示
- 你的第一列数据初始是字符串类型,转为数值型后绘图会更准确,避免matplotlib把它当作分类数据处理。
- 如果需要让离散的有效数据点之间用线条连接,两种方式都可以通过设置
linestyle='-'实现,这样能更直观展示长期趋势。
内容的提问来源于stack exchange,提问作者ZakS
相关产品推荐
相关产品推荐

