如何用三个DataFrame替代Python循环绘制企鹅散点图
解决方案
直接利用你已经创建好的三个独立物种DataFrame(adelie、chinstrap、gentoo),分别调用绘图函数即可替代原有的groupby循环逻辑,最终图表效果和原代码完全一致。以下是修改后的完整代码:
import pandas as pd import matplotlib.pyplot as plt URL= 'https://gist.githubusercontent.com/anibali/c2abc8cab4a2f7b0a6518d11a67c693c/raw/3b1bb5264736bb762584104c9e7a828bef0f6ec8/penguins.csv' df = pd.read_csv(URL) # --- 图表1:体重 vs 喙长 --- fig, ax = plt.subplots() # 提前创建三个物种的独立DataFrame adelie = df[df['species'] == 'Adelie'] chinstrap = df[df['species'] == 'Chinstrap'] gentoo = df[df['species'] == 'Gentoo'] # 分别绘制每个物种的散点 ax.plot(adelie['body_mass_g'], adelie['bill_length_mm'], marker='o', linestyle='', label='Adelie') ax.plot(chinstrap['body_mass_g'], chinstrap['bill_length_mm'], marker='o', linestyle='', label='Chinstrap') ax.plot(gentoo['body_mass_g'], gentoo['bill_length_mm'], marker='o', linestyle='', label='Gentoo') ax.set_title('Penguin measurements by species') ax.set_xlabel('Body mass (g)') ax.set_ylabel('Bill length (mm)') fig.tight_layout() plt.legend() plt.show() # --- 图表2:体重 vs 喙比例 --- df['bill_proportion'] = df['bill_length_mm'] / df['bill_depth_mm'] fig, ax = plt.subplots() # 直接复用已创建的三个物种DF(或重新创建,结果一致) ax.plot(adelie['body_mass_g'], adelie['bill_proportion'], marker='o', linestyle='', label='Adelie') ax.plot(chinstrap['body_mass_g'], chinstrap['bill_proportion'], marker='o', linestyle='', label='Chinstrap') ax.plot(gentoo['body_mass_g'], gentoo['bill_proportion'], marker='o', linestyle='', label='Gentoo') ax.set_title('Penguin proportions by species') ax.set_xlabel('Body mass (g)') ax.set_ylabel('Bill proportion (length/width)') fig.tight_layout() plt.legend() plt.show()
关键改动说明
- 移除了原代码中冗余的
dataDataFrame构建和groupby循环逻辑,直接使用已筛选好的三个物种独立DataFrame - 对每个物种单独调用
ax.plot(),指定对应的x/y列、标记样式和标签,完全复刻原散点图效果 - 保留了原代码中所有的标题、轴标签、图例和布局设置,确保图表视觉效果完全匹配
补充说明
如果更贴合散点图的语义,可以将ax.plot()替换为ax.scatter(),默认标记就是圆形,无需额外设置marker='o',效果完全一致:
ax.scatter(adelie['body_mass_g'], adelie['bill_length_mm'], label='Adelie')
内容的提问来源于stack exchange,提问作者Nicko Toumbas
相关产品推荐
相关产品推荐

