Pandas堆叠面积图无法隐藏NaN值的问题排查与求助
我拥有如下CSV格式的数据集,记录了不同版本(release)在对应日期(Date)的计数(count):
Date,release,count 2019-03-01,buster,0 2019-03-01,jessie,1 2019-03-01,stretch,74 2019-08-15,buster,25 2019-08-15,jessie,1 2019-08-15,stretch,49 2019-10-07,buster,35 2019-10-07,jessie,1 2019-10-07,stretch,43 2019-10-08,buster,40 2019-10-08,jessie,1 2019-10-08,stretch,38 2019-10-09,buster,46 2019-10-09,jessie,1 2019-10-09,stretch,33 2019-10-23,buster,46 2019-10-23,jessie,1 2019-10-23,stretch,31 2019-11-25,buster,46 2019-11-25,jessie,1 2019-11-25,stretch,29 2020-01-13,buster,48 2020-01-13,jessie,1 2020-01-13,stretch,28 2020-01-29,buster,50 2020-01-29,jessie,1 2020-01-29,stretch,26 2020-03-10,buster,54 2020-03-10,jessie,1 2020-03-10,stretch,22 2020-04-14,buster,55 2020-04-14,jessie,0 2020-04-14,stretch,21 2020-05-11,buster,57 2020-05-11,jessie,0 2020-05-11,stretch,17 2020-05-25,buster,61 2020-05-25,jessie,0 2020-05-25,stretch,14 2020-06-10,buster,62 2020-06-10,stretch,12 2020-07-01,buster,69 2020-07-01,stretch,3 2020-10-30,buster,74 2020-10-30,stretch,2 2020-11-18,buster,76 2020-11-18,stretch,2 2021-08-26,bullseye,1 2021-08-26,buster,86 2021-08-26,stretch,1 2021-10-08,bullseye,4 2021-10-08,buster,86 2021-10-08,stretch,1 2021-11-11,bullseye,4 2021-11-11,buster,84 2021-11-11,stretch,1 2021-11-17,bullseye,4 2021-11-17,buster,85 2021-11-17,stretch,0
使用以下Python代码加载数据、透视转换并绘制堆叠面积图:
import pandas as pd import matplotlib.pyplot as plt # Load the data df = pd.read_csv('subset.csv') # Pivot the data to a suitable format for plotting df = df.pivot_table(index="Date", columns='release', values='count', aggfunc='sum') # Convert the index to datetime and sort it df.index = pd.to_datetime(df.index) print(df) # Plotting the data with filled areas fig, ax = plt.subplots(figsize=(12, 6)) df.plot(ax=ax, kind="area", stacked=True) plt.show()
透视后的DataFrame中,jessie版本在2020-05-25之后的值均为NaN,bullseye版本在2021-08-26之前的值均为NaN,但绘制的堆叠面积图中,jessie的线条并未在2020-05-25处终止,而是持续显示至图表末尾;bullseye的线条也从图表起始处就开始显示。
改用kind="bar"绘制堆叠柱状图时无此问题,但柱状图无法正确适配时间轴的缩放特性。
请问为何Pandas/matplotlib会将NaN值当作0绘制而非隐藏?dropna方法无法解决此问题(它会删除整行/列而非单个单元格),该如何处理才能让堆叠面积图正确隐藏NaN值对应的部分?
问题原因
Pandas的area绘图方法在处理堆叠面积图时,会自动对NaN值进行线性插值或者将其视为连续序列的一部分,而不是直接截断。这是因为面积图需要连续的填充区域,默认的绘图逻辑会尝试补全缺失值,导致NaN的位置被填充为前后值的延续,看起来像是当作0处理,但实际是插值后的结果。而柱状图是离散的,每个日期独立绘制,所以不会有这个问题。
处理方法
要让堆叠面积图正确隐藏NaN对应的部分,需要手动处理每个版本的时间序列,只保留有数据的区间,然后分别绘制每个版本的面积图,再手动堆叠:
- 遍历每个版本(列),提取该版本非NaN值对应的日期和计数,确保时间序列连续且仅包含有效数据。
- 对每个版本的有效数据,计算堆叠的基准值(即前面所有版本的累计和)。
- 使用matplotlib的
fill_between函数手动绘制每个版本的填充区域。
修改后的代码如下:
import pandas as pd import matplotlib.pyplot as plt # Load the data df = pd.read_csv('subset.csv') # Pivot the data to a suitable format for plotting df_pivot = df.pivot_table(index="Date", columns='release', values='count', aggfunc='sum') # Convert the index to datetime and sort it df_pivot.index = pd.to_datetime(df_pivot.index) df_pivot = df_pivot.sort_index() fig, ax = plt.subplots(figsize=(12, 6)) # 初始化堆叠基准值 bottom = pd.Series([0]*len(df_pivot.index), index=df_pivot.index) # 遍历每个版本,只绘制有有效数据的部分 for release in df_pivot.columns: # 获取当前版本的非NaN数据 data = df_pivot[release].dropna() if not data.empty: # 对齐基准值到当前版本的时间区间 aligned_bottom = bottom.loc[data.index] # 绘制填充区域 ax.fill_between(data.index, aligned_bottom, aligned_bottom + data, label=release) # 更新基准值(仅在有数据的区间更新) bottom.loc[data.index] += data # 设置图表属性 ax.set_xlabel('Date') ax.set_ylabel('Count') ax.set_title('Stacked Area Plot of Release Counts') ax.legend() plt.xticks(rotation=45) plt.tight_layout() plt.show()
代码说明
- 遍历每个版本时,用
dropna()提取该版本有数据的日期和计数,确保只处理有效区间。 bottom变量记录当前堆叠的基准高度,每个版本的填充区域从bottom开始,到bottom+data结束。- 仅在当前版本有数据的日期区间更新
bottom,这样后续版本的堆叠只会在有效区间进行,NaN对应的区域不会被填充。
这种方法完全避免了Pandas自动插值的问题,确保每个版本的线条只在有数据的区间显示,完美符合需求。
内容的提问来源于stack exchange,提问作者anarcat

