如何用Pandas为大量数据生成清晰的EASTVEL-Z_NAP折线图?
解决方案
针对数据量过大导致折线图杂乱的问题,你可以通过以下几种方式生成清晰的EASTVEL随Z_NAP变化的趋势折线:
1. 按水深(Z_NAP)分组聚合
同一水深往往对应多个流速数据,通过对每个Z_NAP取值计算均值/中位数/百分位数,将多对一的关系转为一对一,大幅减少数据点数量:
import pandas as pd import matplotlib.pyplot as plt # 按Z_NAP分组,计算EASTVEL的均值(也可替换为median()、quantile(0.5)) aggregated = frame_3_LW.groupby('Z_NAP')['EASTVEL'].mean().reset_index() # 绘制趋势折线图 plt.figure(figsize=(10,6)) plt.plot(aggregated['EASTVEL'], aggregated['Z_NAP'], linewidth=2) plt.xlabel('EASTVEL (Eastward Velocity)') plt.ylabel('Z_NAP (Water Depth)') plt.title('EASTVEL vs Z_NAP Trend') plt.grid(True, alpha=0.3) plt.show()
若Z_NAP是连续浮点值,可先对其分箱(按固定间隔划分水深区间)再聚合:
# 按0.5米间隔对Z_NAP分箱 frame_3_LW['Z_NAP_bin'] = pd.cut(frame_3_LW['Z_NAP'], bins=range(int(frame_3_LW['Z_NAP'].min()), int(frame_3_LW['Z_NAP'].max())+1, 0.5)) # 按箱分组计算EASTVEL均值 aggregated_bin = frame_3_LW.groupby('Z_NAP_bin')['EASTVEL'].mean().reset_index() # 将区间转为中间值用于绘图 aggregated_bin['Z_NAP_mid'] = aggregated_bin['Z_NAP_bin'].apply(lambda x: x.mid) plt.figure(figsize=(10,6)) plt.plot(aggregated_bin['EASTVEL'], aggregated_bin['Z_NAP_mid'], linewidth=2) plt.xlabel('EASTVEL (Eastward Velocity)') plt.ylabel('Z_NAP (Water Depth)') plt.title('EASTVEL vs Z_NAP Trend (Binned Depth)') plt.grid(True, alpha=0.3) plt.show()
2. 数据平滑处理
若不想丢失原始数据的分布特征,可通过滚动窗口平滑或LOESS局部回归拟合生成趋势线:
滚动窗口平滑
先按Z_NAP排序,再用滚动窗口计算均值弱化噪声:
# 按Z_NAP排序数据 sorted_df = frame_3_LW.sort_values('Z_NAP') # 滚动窗口平滑,窗口大小可根据数据密度调整(示例为50个数据点) sorted_df['EASTVEL_smoothed'] = sorted_df['EASTVEL'].rolling(window=50, center=True).mean() plt.figure(figsize=(10,6)) # 可选:绘制原始散点(调低透明度避免遮挡) plt.scatter(sorted_df['EASTVEL'], sorted_df['Z_NAP'], alpha=0.1, label='Raw Data') # 绘制平滑后的趋势折线 plt.plot(sorted_df['EASTVEL_smoothed'], sorted_df['Z_NAP'], linewidth=2, color='red', label='Smoothed Trend') plt.xlabel('EASTVEL (Eastward Velocity)') plt.ylabel('Z_NAP (Water Depth)') plt.title('EASTVEL vs Z_NAP with Rolling Smoothing') plt.legend() plt.grid(True, alpha=0.3) plt.show()
LOESS局部回归拟合
适合捕捉非线性趋势,用statsmodels实现:
import statsmodels.api as sm # 按Z_NAP排序数据 sorted_df = frame_3_LW.sort_values('Z_NAP') # 拟合LOESS,frac参数控制平滑程度(0-1,值越大曲线越平滑) lowess = sm.nonparametric.lowess(sorted_df['EASTVEL'], sorted_df['Z_NAP'], frac=0.1) # 提取拟合后的结果 smoothed_x = lowess[:, 1] smoothed_y = lowess[:, 0] plt.figure(figsize=(10,6)) plt.scatter(sorted_df['EASTVEL'], sorted_df['Z_NAP'], alpha=0.1, label='Raw Data') plt.plot(smoothed_x, smoothed_y, linewidth=2, color='red', label='LOESS Smoothed Trend') plt.xlabel('EASTVEL (Eastward Velocity)') plt.ylabel('Z_NAP (Water Depth)') plt.title('EASTVEL vs Z_NAP with LOESS Smoothing') plt.legend() plt.grid(True, alpha=0.3) plt.show()
3. 降采样数据
若数据是高频率采样,可按固定间隔抽取数据点:
# 按Z_NAP排序后,每隔10个数据点取一个 sampled_df = frame_3_LW.sort_values('Z_NAP').iloc[::10, :] plt.figure(figsize=(10,6)) plt.plot(sampled_df['EASTVEL'], sampled_df['Z_NAP'], linewidth=2) plt.xlabel('EASTVEL (Eastward Velocity)') plt.ylabel('Z_NAP (Water Depth)') plt.title('EASTVEL vs Z_NAP (Downsampled Data)') plt.grid(True, alpha=0.3) plt.show()
内容的提问来源于stack exchange,提问作者Emile Everduim
相关产品推荐
相关产品推荐

