如何绘制序列事件数据value字段随时间变化的缺失值(NA)趋势图
缺失值随时间变化趋势绘制方案
核心逻辑
- 按
time字段对数据分组,统计每个时间节点下value列的NA数量或占比 - 以时间为横轴,NA统计值为纵轴绘制折线图,即可直观呈现缺失值随时间的变化趋势
- 你提供的示例数据统计结果为:time=1时NA数为0,time=2为1,time=3为2,time=4为3,time=5为5,完全符合随时间递增的特征
Python 实现(pandas + matplotlib)
import pandas as pd import matplotlib.pyplot as plt # 替换为你的数据读取逻辑,支持csv、excel等多种格式 df = pd.read_csv("your_data_path.csv") # 按time分组统计value列的NA数量,如需统计缺失率将sum改为mean再乘以100即可 na_stats = df.groupby("time")["value"].apply(lambda x: x.isna().sum()).reset_index(name="na_count") # 绘制趋势图 plt.figure(figsize=(8, 4)) plt.plot(na_stats["time"], na_stats["na_count"], marker="o", color="#2c7fb8", linewidth=2) plt.xlabel("时间") plt.ylabel("value字段NA数量") plt.title("value字段缺失数随时间变化趋势") plt.grid(axis="y", linestyle="--", alpha=0.7) plt.show()
R 实现(dplyr + ggplot2)
library(dplyr) library(ggplot2) # 替换为你的数据读取逻辑,na.strings指定将文本NA识别为缺失值 df <- read.csv("your_data_path.csv", na.strings = "NA") # 按time分组统计NA数量,如需缺失率将sum改为mean再乘以100即可 na_stats <- df %>% group_by(time) %>% summarise(na_count = sum(is.na(value))) # 绘制趋势图 ggplot(na_stats, aes(x = time, y = na_count)) + geom_line(color = "#2c7fb8", linewidth = 1.2) + geom_point(size = 3) + labs(x = "时间", y = "value字段NA数量", title = "value字段缺失数随时间变化趋势") + theme_bw() + theme(panel.grid.minor = element_blank())
内容的提问来源于stack exchange,提问作者cliu
相关产品推荐
相关产品推荐

