You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何分析千万级DataFrame中时间与速度的关联及年度变化?

分析时间与速度关联的解决方案

针对1200万行量级的数据集,直接绘制回归图因数据点过度重叠导致可视化失效,以下是分步骤的高效分析方案:

1. 数据预处理

首先完成时间特征的提取,为后续分组分析做准备:

  • 确保时间列是标准datetime格式:df['timestamp'] = pd.to_datetime(df['timestamp'])
  • 提取核心特征:新增hour(0-23,代表一天中的小时)和year列,用于分组聚合
import pandas as pd
import numpy as np

# 假设时间列名为'timestamp',速度列名为'speed'
df['timestamp'] = pd.to_datetime(df['timestamp'])
df['hour'] = df['timestamp'].dt.hour
df['year'] = df['timestamp'].dt.year

2. 数据聚合(解决百万级数据可视化痛点)

直接绘制原始数据会导致视觉混乱,必须先按年份+小时维度聚合,计算统计量消除噪声:

# 聚合计算每个时段的速度均值、中位数及置信区间
agg_df = df.groupby(['year', 'hour'])['speed'].agg(
    mean_speed='mean',
    median_speed='median',
    std_speed='std',
    sample_count='count'
).reset_index()

# 计算95%置信区间(用于展示数据波动的统计显著性)
agg_df['ci_low'] = agg_df['mean_speed'] - 1.96 * agg_df['std_speed'] / np.sqrt(agg_df['sample_count'])
agg_df['ci_high'] = agg_df['mean_speed'] + 1.96 * agg_df['std_speed'] / np.sqrt(agg_df['sample_count'])

3. 针对性可视化方案

3.1 分年份的小时-速度趋势折线图

清晰对比各年份的时段速度变化,同时用置信区间填充展示数据稳定性:

import seaborn as sns
import matplotlib.pyplot as plt

plt.figure(figsize=(12,6))
sns.lineplot(data=agg_df, x='hour', y='mean_speed', hue='year', marker='o', linewidth=2)

# 为每个年份添加置信区间填充
for year in agg_df['year'].unique():
    year_subset = agg_df[agg_df['year'] == year]
    plt.fill_between(year_subset['hour'], year_subset['ci_low'], year_subset['ci_high'], alpha=0.2)

plt.xlabel('一天中的小时')
plt.ylabel('平均速度')
plt.title('各年份小时-速度趋势对比')
plt.xticks(range(0,24))
plt.legend(title='年份')
plt.show()

3.2 年份-小时速度热力图

快速定位不同年份的最佳连接时段,直观呈现二维维度的速度差异:

# 转换为透视表格式适配热力图
heatmap_data = agg_df.pivot(index='year', columns='hour', values='mean_speed')

plt.figure(figsize=(14,6))
sns.heatmap(heatmap_data, annot=True, fmt='.1f', cmap='viridis', cbar=True)
plt.xlabel('一天中的小时')
plt.ylabel('年份')
plt.title('年份-小时速度热力图')
plt.show()

3.3 小时速度分布箱线图

查看每个小时的速度离散程度,对比不同年份的波动差异:

plt.figure(figsize=(12,6))
sns.boxplot(data=df, x='hour', y='speed', hue='year')
plt.xlabel('一天中的小时')
plt.ylabel('速度')
plt.title('各年份小时速度分布箱线图')
plt.legend(title='年份')
plt.show()

4. 统计验证关联变化

通过统计模型量化年份对小时-速度关联的影响:

  • 加入交互项的线性回归:验证hour与year的交互项是否显著,判断年份是否改变时段对速度的影响
import statsmodels.api as sm

df['hour_year_interaction'] = df['hour'] * df['year']
X = sm.add_constant(df[['hour', 'year', 'hour_year_interaction']])
y = df['speed']

# 拟合模型(若数据量过大,可抽样后拟合)
model = sm.OLS(y.sample(frac=0.1), X.sample(frac=0.1)).fit()
print(model.summary())

若hour_year_interaction的p值<0.05,说明年份显著改变了小时与速度的关联关系。

性能优化提示

  • 处理1200万行数据时,优先使用pandas矢量化操作,避免循环
  • 若内存不足,可采用分块读取(pd.read_csv(chunksize=100000))或使用Dask库进行分布式处理
  • 统计模型拟合时可抽样10%-20%数据,在保证结果准确性的前提下提升速度

内容的提问来源于stack exchange,提问作者dim GOROH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 03:15:38