Python中如何基于DataFrame数据在同一面板绘制重叠直方图
解决DataFrame绘制重叠直方图的问题
没问题,我来帮你搞定这个问题~你之前的代码出错是因为matplotlib的plt.hist()不支持直接传入DataFrame对象并指定y参数,你需要先从DataFrame中提取出目标列的数据,再传入直方图函数。下面给你两种可行的解决方案:
方法1:提取DataFrame列数据,用matplotlib原生绘制
直接从你的cv1_weight和cv2_weight中取出AGW列,再传入plt.hist(),同时可以加上透明度(alpha)和标签(label)让重叠的直方图更清晰:
import numpy as np import pandas as pd import matplotlib.pyplot as plt # 生成原始数据 cv1 = np.random.normal(50, 4, 1000) cv2 = np.random.normal(40, 7, 1000) # 构造DataFrame cv1_weight = pd.DataFrame({"Genotype": ["cv1"]*1000,"AGW": cv1}) cv2_weight = pd.DataFrame({"Genotype": ["cv2"]*1000,"AGW": cv2}) # 绘制重叠直方图 plt.hist(cv1_weight["AGW"], bins=np.linspace(0, 70, 100), alpha=0.5, label="cv1") plt.hist(cv2_weight["AGW"], bins=np.linspace(0, 70, 100), alpha=0.5, label="cv2") plt.xlabel("Grain weight (mg)", size=12) plt.ylabel("Frequency_Histogram", size=12) plt.legend() # 显示图例区分两组数据 plt.show()
方法2:用Pandas内置的绘图接口(更简洁)
Pandas的DataFrame自带plot.hist()方法,你可以先把两个DataFrame合并,再一次性绘制分组直方图,代码更简洁:
import numpy as np import pandas as pd import matplotlib.pyplot as plt # 生成数据并合并DataFrame cv1 = np.random.normal(50, 4, 1000) cv2 = np.random.normal(40, 7, 1000) df = pd.concat([ pd.DataFrame({"Genotype": ["cv1"]*1000,"AGW": cv1}), pd.DataFrame({"Genotype": ["cv2"]*1000,"AGW": cv2}) ]) # 绘制重叠直方图 df.plot.hist(column="AGW", by="Genotype", bins=np.linspace(0, 70, 100), alpha=0.5, figsize=(8,5)) plt.xlabel("Grain weight (mg)", size=12) plt.ylabel("Frequency_Histogram", size=12) plt.title("") # 去掉自动生成的冗余标题 plt.show()
两种方法都能实现你想要的同一面板内的重叠直方图效果,方法2更适合处理结构化的DataFrame数据,推荐尝试~
内容的提问来源于stack exchange,提问作者J.K Kim
相关产品推荐
相关产品推荐

