You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何筛选数据集:排除practice地点30天内重复的unique_ID记录

数据集筛选方案

核心筛选逻辑

保留满足以下任一条件的记录:

  • 检测地点location不是"practice";
  • 检测地点是"practice",但同一unique_ID的所有检测记录间隔均超过30天(无30天内重复的情况)。

R语言实现(基于dplyr包)

library(dplyr)
library(lubridate)

# 转换日期列为标准格式
df <- df %>%
  mutate(test_date = ymd(test_date)) %>%
  # 按用户ID和检测地点分组
  group_by(unique_ID, location) %>%
  # 标记当前组是否存在30天内的重复记录
  mutate(has_recent_dup = any(abs(difftime(test_date, lag(test_date), units = "days")) <= 30, na.rm = TRUE)) %>%
  ungroup() %>%
  # 筛选符合要求的记录
  filter(location != "practice" | (location == "practice" & !has_recent_dup))

Python语言实现(基于pandas包)

import pandas as pd

# 转换日期列格式
df['test_date'] = pd.to_datetime(df['test_date'])

# 计算同组内每条记录与上一条的天数差
df['days_diff'] = df.groupby(['unique_ID', 'location'])['test_date'].diff().dt.days.abs()

# 标记组内是否存在30天内的重复
df['has_recent_dup'] = df.groupby(['unique_ID', 'location'])['days_diff'].transform(lambda x: any(x <= 30))

# 筛选记录并清理临时列
filtered_df = df[(df['location'] != 'practice') | ((df['location'] == 'practice') & (~df['has_recent_dup']))]
filtered_df = filtered_df.drop(columns=['days_diff', 'has_recent_dup'])

内容的提问来源于stack exchange,提问作者Thomas Rosengren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 16:04:58