如何为含重复分组值的CSV数据绘制折线图
如何为含重复分类列的CSV数据绘制折线图

我有一个CSV文件,其中某一列的值存在重复,且对应有数值数据(如示例中的location列与new_cases列)。请问如何为这些数据绘制折线图?
我尝试了以下代码,但运行失败:
import matplotlib.pyplot as plt import pandas as pd data = {'location': ['Afghanistan'] * 5 + ['Africa'] * 4, 'new_cases': [3, 0, 0, 3, 6, 0, 1, 0, 0]} newData = pd.DataFrame(data) fig, ax = plt.subplots(figsize=(15,7)) byLoc = newData.groupby('location').count()['new_cases'].unstack().plot(ax=ax)
报错信息
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In [141], line 2 1 fig, ax = plt.subplots(figsize=(15,7)) ----> 2 byLoc = newData.groupby('location').count()['new_cases'].unstack().plot(ax=ax) File ~\anaconda3\envs\py11\Lib\site-packages\pandas\core\series.py:4455, in Series.unstack(self, level, fill_value) 4412 """ 4413 Unstack, also known as pivot, Series with MultiIndex to produce DataFrame. 4414 (...) 4451 b 2 4 4452 """ 4453 from pandas.core.reshape.reshape import unstack -> 4455 return unstack(self, level, fill_value) File ~\anaconda3\envs\py11\Lib\site-packages\pandas\core\reshape\reshape.py:483, in unstack(obj, level, fill_value) 478 return obj.T.stack(dropna=False) 479 elif not isinstance(obj.index, MultiIndex): 480 # GH 36113 481 # Give nicer error messages when unstack a Series whose 482 # Index is not a MultiIndex. -> 483 raise ValueError( 484 f"index must be a MultiIndex to unstack, {type(obj.index)} was passed" 485 ) 486 else: 487 if is_1d_only_ea_dtype(obj.dtype): ValueError: unstack操作要求索引为MultiIndex,但传入的是<class 'pandas.core.indexes.base.Index'>类型
错误原因
你的代码有两个问题:
groupby('location').count()统计的是每个分类的行数,并非你需要的new_cases时序数据;unstack()仅适用于**多级索引(MultiIndex)**对象,而groupby('location')后得到的是普通索引的Series,因此触发报错。
正确实现方式
折线图需要展示每个分类下new_cases随序列(或时间)的变化,以下是两种简单可行的方法:
方法1:分组循环绘制
import matplotlib.pyplot as plt import pandas as pd data = {'location': ['Afghanistan'] * 5 + ['Africa'] * 4, 'new_cases': [3, 0, 0, 3, 6, 0, 1, 0, 0]} newData = pd.DataFrame(data) fig, ax = plt.subplots(figsize=(15,7)) # 按location分组,为每个组单独绘制折线 for name, group in newData.groupby('location'): ax.plot(group.index, group['new_cases'], label=name) ax.set_xlabel('数据序号') ax.set_ylabel('新增病例数') ax.set_title('不同地区新增病例变化') ax.legend() plt.show()
方法2:利用pandas分组直接绘图
pandas的groupby对象支持直接调用plot,自动生成多分类折线:
import matplotlib.pyplot as plt import pandas as pd data = {'location': ['Afghanistan'] * 5 + ['Africa'] * 4, 'new_cases': [3, 0, 0, 3, 6, 0, 1, 0, 0]} newData = pd.DataFrame(data) fig, ax = plt.subplots(figsize=(15,7)) # 分组后直接绘制折线图 newData.groupby('location')['new_cases'].plot(ax=ax, legend=True) ax.set_xlabel('数据序号') ax.set_ylabel('新增病例数') ax.set_title('不同地区新增病例变化') plt.show()
补充:含时间轴的场景
如果你的实际数据包含时间列(如date),只需将x轴替换为时间字段即可:
import matplotlib.pyplot as plt import pandas as pd # 模拟带日期的数据 data = {'location': ['Afghanistan']*5 + ['Africa']*4, 'date': pd.date_range(start='2020-01-01', periods=5).tolist() + pd.date_range(start='2020-01-01', periods=4).tolist(), 'new_cases': [3,0,0,3,6,0,1,0,0]} newData = pd.DataFrame(data) fig, ax = plt.subplots(figsize=(15,7)) for name, group in newData.groupby('location'): ax.plot(group['date'], group['new_cases'], label=name) ax.set_xlabel('日期') ax.set_ylabel('新增病例数') ax.set_title('不同地区新增病例变化') ax.legend() plt.xticks(rotation=45) plt.show()
内容的提问来源于stack exchange,提问作者Elio Baharan
相关产品推荐
相关产品推荐

