You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中按年龄范围设置散点图数据点的颜色?

按年龄分组设置散点图颜色的解决方案

针对你的需求,这里提供两种简单易上手的实现方法,基于Python的pandas和matplotlib库:

方法一:添加年龄分组列后分组绘图

逻辑清晰,适合新手理解:

import pandas as pd
import matplotlib.pyplot as plt

# 加载高尔夫数据集(替换为你的数据集路径)
# df = pd.read_csv("golf_dataset.csv")

# 1. 给数据添加年龄分组标签
def get_age_group(age):
    if 20 <= age <=29:
        return "20-29岁"
    elif 30 <= age <=39:
        return "30-39岁"
    elif 40 <= age <=49:
        return "40-49岁"
    else:
        return "其他"

df['age_group'] = df['AGE'].apply(get_age_group)

# 2. 定义分组对应颜色
color_map = {
    "20-29岁": "blue",
    "30-39岁": "red",
    "40-49岁": "yellow"
}

# 3. 绘制分组散点图
plt.figure(figsize=(10,6))
for group, color in color_map.items():
    group_data = df[df['age_group'] == group]
    plt.scatter(group_data['Fairway Hit%'], group_data['Average Driving Distance'], 
                color=color, label=group, alpha=0.6)

# 添加图表信息
plt.xlabel("Fairway Hit%")
plt.ylabel("Average Driving Distance")
plt.title("开球距离 vs 球道上球率(按年龄分组)")
plt.legend()
plt.show()

方法二:直接生成颜色列表(无需额外列)

如果不想新增数据列,可用numpy的where函数直接生成每个点的颜色:

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

# 加载数据集
# df = pd.read_csv("golf_dataset.csv")

# 生成每个数据点的颜色
colors = np.where((df['AGE'] >=20) & (df['AGE'] <=29), 'blue',
                  np.where((df['AGE'] >=30) & (df['AGE'] <=39), 'red',
                           np.where((df['AGE'] >=40) & (df['AGE'] <=49), 'yellow', 'gray')))

# 绘制散点图
plt.figure(figsize=(10,6))
plt.scatter(df['Fairway Hit%'], df['Average Driving Distance'], color=colors, alpha=0.6)

# 手动添加图例
from matplotlib.patches import Patch
legend_elements = [Patch(facecolor='blue', label='20-29岁'),
                   Patch(facecolor='red', label='30-39岁'),
                   Patch(facecolor='yellow', label='40-49岁')]
plt.legend(handles=legend_elements)

plt.xlabel("Fairway Hit%")
plt.ylabel("Average Driving Distance")
plt.title("开球距离 vs 球道上球率(按年龄分组)")
plt.show()

你之前代码出错的原因

你尝试写'AGE' >=20 <=29这种条件是错误的,在Python中对Series做范围判断时,需要用括号包裹单个条件,并用&连接逻辑与,正确写法是(df['AGE'] >=20) & (df['AGE'] <=29)。

内容的提问来源于stack exchange,提问作者8ironanalytics

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 00:45:14