You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr或aggregate函数按10分钟间隔计算几何均值?

Calculating 10-Minute Interval Geometric Means for Your Time Series Data

Got it, let's work through this problem step by step. Calculating geometric means in 10-minute time bins is totally manageable with either Python (Pandas) or R—here's how to handle both, since I don't know which tool you prefer.

Using Python (Pandas)

First, you need to make sure your TimeDate column is recognized as datetime data—Pandas relies on that to group by time intervals. Then, grouping into 10-minute chunks and computing the geometric mean is straightforward.

Step 1: Load and prepare your data

import pandas as pd
import scipy.stats as stats

# Load your data (replace with your actual file path or data source)
df = pd.read_csv('your_data.csv', parse_dates=['TimeDate'])

# Optional: Set TimeDate as the index to simplify grouping
df = df.set_index('TimeDate')

Step 2: Group into 10-minute intervals and compute geometric mean

Pandas doesn't have a built-in geometric mean function, so we'll use scipy.stats.gmean for this. We'll apply it to the columns you care about (like diam or ratio):

# Define the 10-minute interval frequency
freq = '10min'

# Group by intervals and calculate geometric means for target columns
interval_gmean = df.groupby(pd.Grouper(freq=freq)).agg(
    diam_gmean=('diam', stats.gmean),
    ratio_gmean=('ratio', stats.gmean)
).dropna()  # Remove intervals with no observations (omit if you want to keep NaNs)

# Preview the result
print(interval_gmean.head())

Note: Geometric mean only makes sense for positive values—double-check that your diam and ratio columns don't have zeros or negatives before running this.


Using R

If you're working in R, lubridate handles time parsing easily, and dplyr simplifies grouping. The psych package has a ready-to-use geometric.mean function, or you can calculate it manually using logarithms.

Step 1: Load libraries and prepare data

library(dplyr)
library(lubridate)
library(psych)

# Load your data (replace with your file path)
df <- read.csv("your_data.csv")

# Parse TimeDate as datetime
df$TimeDate <- ymd_hms(df$TimeDate)

# Create 10-minute interval bins
df$interval <- floor_date(df$TimeDate, unit = "10 minutes")

Step 2: Compute geometric mean per interval

# Group by interval and calculate geometric means
interval_gmean <- df %>%
  group_by(interval) %>%
  summarise(
    diam_gmean = geometric.mean(diam, na.rm = TRUE),
    ratio_gmean = geometric.mean(ratio, na.rm = TRUE)
  ) %>%
  filter(!is.na(diam_gmean))  # Drop empty intervals (optional)

# View the result
head(interval_gmean)

Alternative manual calculation (no need for psych package):

diam_gmean = exp(mean(log(diam), na.rm = TRUE))

This works the same way, as long as all values are positive.


Quick Tips

  • By default, intervals are aligned to the hour (e.g., 8:20-8:30, 8:30-8:40). If you want bins to start from your first observation time instead, let me know—I can adjust the code for that!
  • If you have missing values in your data, the na.rm = TRUE parameter ensures they don't break the calculation.

内容的提问来源于stack exchange,提问作者user2928318

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 02:23:26