如何从CSV文件的Pandas DataFrame中查找局部最大值并标注到图表?
从Pandas DataFrame查找局部最大值并标注到图表
1. 数据准备(读取CSV或导入已有数据)
如果是从CSV读取数据,直接用pd.read_csv()即可;针对你提供的示例数据,可按如下方式构建DataFrame:
import pandas as pd import matplotlib.pyplot as plt from scipy.signal import find_peaks # 你的x、y数据 x = [1598.78, 1596.85, 1594.92, 1592.99, 1591.07, 1589.14, 1587.21, 1585.28, 1583.35, 1581.42, 1579.49, 1577.57, 1575.64, 1573.71, 1571.78, 1569.85, 1567.92, 1565.99, 1564.07, 1562.14, 1560.21, 1558.28, 1556.35, 1554.42, 1552.49, 1550.57, 1548.64, 1546.71, 1544.78, 1542.85, 1540.92, 1538.99, 1537.07, 1535.14, 1533.21, 1531.28, 1529.35, 1527.42, 1525.49, 1523.57, 1521.64, 1519.71, 1517.78, 1515.85, 1513.92, 1511.99, 1510.07, 1508.14, 1506.21, 1504.28, 1502.35, 1500.42, 1498.49, 1496.57, 1494.64, 1492.71, 1490.78, 1488.85, 1486.92, 1484.99, 1483.07, 1481.14, 1479.21, 1477.28, 1475.35, 1473.42, 1471.49, 1469.57, 1467.64, 1465.71, 1463.78, 1461.85, 1459.92, 1457.99, 1456.07, 1454.14, 1452.21, 1450.28, 1448.35, 1446.42, 1444.49, 1442.57, 1440.64, 1438.71, 1436.78, 1434.85, 1432.92, 1430.99, 1429.07, 1427.14, 1425.21, 1423.28, 1421.35, 1419.42, 1417.49, 1415.57, 1413.64, 1411.71, 1409.78, 1407.85, 1405.92, 1403.99, 1402.07, 1400.14] y = [0.640, 0.624, 0.609, 0.594, 0.581, 0.569, 0.558, 0.547, 0.537, 0.530, 0.523, 0.516, 0.508, 0.502, 0.497, 0.491, 0.487, 0.484, 0.481, 0.480, 0.479, 0.482, 0.490, 0.503, 0.520, 0.542, 0.566, 0.586, 0.600, 0.606, 0.593, 0.569, 0.557, 0.548, 0.538, 0.531, 0.527, 0.524, 0.522, 0.522, 0.523, 0.525, 0.526, 0.527, 0.530, 0.534, 0.536, 0.539, 0.547, 0.553, 0.557, 0.563, 0.573, 0.599, 0.654, 0.738, 0.852, 0.891, 0.810, 0.744, 0.711, 0.694, 0.686, 0.683, 0.683, 0.690, 0.700, 0.706, 0.713, 0.723, 0.731, 0.732, 0.737, 0.756, 0.779, 0.786, 0.790, 0.794, 0.802, 0.815, 0.827, 0.832, 0.831, 0.826, 0.823, 0.828, 0.834, 0.834, 0.832, 0.832, 0.831, 0.825, 0.816, 0.804, 0.798, 0.794, 0.786, 0.775, 0.764, 0.752, 0.739, 0.722, 0.708, 0.697] df = pd.DataFrame({"x": x, "y": y})
2. 检测局部最大值
用scipy.signal.find_peaks精准检测峰值,可通过参数过滤无效小峰:
# 查找峰值:height设置最小高度阈值,distance设置峰值间最小数据点间隔 peaks, props = find_peaks(df["y"], height=0.5, distance=10) # 提取峰值对应的x、y数据 peak_points = df.iloc[peaks]
3. 绘制曲线并标注峰值
用Matplotlib绘制原始曲线,再通过plt.text在峰值位置标注数值:
plt.figure(figsize=(12, 6)) plt.plot(df["x"], df["y"], color="#1f77b4", linewidth=1.5, label="数据曲线") # 逐个标注峰值的y值,调整偏移量避免与曲线重叠 for _, peak in peak_points.iterrows(): plt.text(peak["x"], peak["y"] + 0.015, f"{peak['y']:.3f}", ha="center", va="bottom", fontsize=10, color="#ff7f0e") # 图表基础设置 plt.xlabel("X轴", fontsize=12) plt.ylabel("Y轴", fontsize=12) plt.title("带局部最大值标注的曲线", fontsize=14) plt.legend() plt.grid(alpha=0.3) plt.show()
参数调整说明
find_peaks的height:过滤低于该值的小波动,避免误判;find_peaks的distance:控制两个峰值的最小间隔,避免相邻近峰被识别;plt.text的偏移量:通过peak["y"] + 0.015调整文字在峰值上方的位置,可根据图表比例修改数值。
内容的提问来源于stack exchange,提问作者cbornes
相关产品推荐
相关产品推荐

