You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何缩放数据点以优化Matplotlib大数据绘图效果?

解决Matplotlib处理大数值数据时绘图元素拥挤的问题

你的问题核心是每组内两个X值的差值相对于全局X轴的总范围来说太小,比如第一组两个X值仅相差约14万,但全局X轴跨度超过5500万,导致连线被压缩得几乎不可见,散点也挤在一起。以下是几种实用的解决办法:

方法1:为每组数据创建独立子图(共享Y轴)

通过给每组数据单独分配子图,让每个子图只聚焦当前组的X范围,能清晰展示散点和连线:

import matplotlib.pyplot as plt

data_pairs = [
    (24193550, 24335121),
    (42850956, 42993424),
    (45606871, 45749886),
    (60595038, 60738084),
    (5026097, 5170030),
]
y_coordinates = [5, 10, 15, 20, 25]

# 创建横向排列的5个子图,共享Y轴
fig, axes = plt.subplots(nrows=1, ncols=5, figsize=(15, 5), sharey=True)

for ax, (x1, x2), y in zip(axes, data_pairs, y_coordinates):
    ax.scatter(x1, y, color='blue', marker='o', s=50)
    ax.scatter(x2, y, color='blue', marker='o', s=50)
    ax.plot([x1, x2], [y, y], color='red', linestyle='-', linewidth=2)
    # 给每个子图的X轴添加少量边距,让线条更显眼
    margin = (x2 - x1) * 0.1
    ax.set_xlim(x1 - margin, x2 + margin)
    ax.set_title(f"Group {i+1}")

# 添加全局轴标签
fig.text(0.5, 0.04, 'X-Axis', ha='center', fontsize=12)
fig.text(0.04, 0.5, 'Y-Axis', va='center', rotation='vertical', fontsize=12)
plt.tight_layout()
plt.show()

方法2:将X值转换为组内相对偏移量

把每组的第一个X值作为基准点,计算第二个X值的相对差值,这样X轴直接展示每组的差值大小,更直观:

import matplotlib.pyplot as plt

data_pairs = [
    (24193550, 24335121),
    (42850956, 42993424),
    (45606871, 45749886),
    (60595038, 60738084),
    (5026097, 5170030),
]
y_coordinates = [5, 10, 15, 20, 25]

plt.figure(figsize=(10, 5))

for i, (x1, x2) in enumerate(data_pairs):
    y = y_coordinates[i]
    # 计算相对偏移,以本组第一个X值为0点
    rel_x1 = 0
    rel_x2 = x2 - x1
    plt.scatter(rel_x1, y, color='blue', marker='o', s=50)
    plt.scatter(rel_x2, y, color='blue', marker='o', s=50)
    plt.plot([rel_x1, rel_x2], [y, y], color='red', linestyle='-', linewidth=2)
    # 标注每组的原始X范围
    plt.text((rel_x1 + rel_x2)/2, y + 0.5, f"{x1:,} → {x2:,}", ha='center', fontsize=9)

plt.xlabel('Relative X Offset (from first value in group)')
plt.ylabel('Y-Axis')
plt.grid(True)
plt.show()

方法3:在单图中添加Y轴微小偏移避免重叠

如果必须在同一张图展示所有数据,可以给每组的Y值添加微小交替偏移,同时标注差值,既保留全局视角,又能看清每组的连线:

import matplotlib.pyplot as plt

data_pairs = [
    (24193550, 24335121),
    (42850956, 42993424),
    (45606871, 45749886),
    (60595038, 60738084),
    (5026097, 5170030),
]
y_coordinates = [5, 10, 15, 20, 25]
y_offset = 0.3  # 每组Y值的微小偏移量

plt.figure(figsize=(12, 6))

for i, (x1, x2) in enumerate(data_pairs):
    # 交替添加正负偏移,避免相邻组的线条重叠
    y = y_coordinates[i] + (-1)**i * y_offset
    plt.scatter(x1, y, color='blue', marker='o', s=50)
    plt.scatter(x2, y, color='blue', marker='o', s=50)
    plt.plot([x1, x2], [y, y], color='red', linestyle='-', linewidth=2)
    # 标注每组的X差值
    plt.text((x1 + x2)/2, y + 0.5, f"ΔX={x2-x1:,}", ha='center', fontsize=9)

plt.xlabel('X-Axis')
plt.ylabel('Y-Axis')
plt.yticks(y_coordinates)  # 保持原始Y轴刻度不变
plt.grid(True)
plt.show()

内容的提问来源于stack exchange,提问作者BlackCrystal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 10:26:12