You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于pandas实现WhatsApp群聊时间窗口参与者提取,用于Random Forest回复者预测

Solution for Adding Time-Window Participant Features to WhatsApp Chat Data

Got it, let's tackle this problem step by step. You're working on predicting message responders in an 11-person WhatsApp group using Random Forest, and need to add features tracking which members participated in the conversation within a specified time window (15 minutes in your example). Below are implementations for both of your desired feature formats.

First: Single Column with Participant List

First off, make sure your date_time column is properly converted to a datetime type—this is critical for accurate time window calculations. Then we'll use a row-wise check to pull all unique members who sent messages in the 15 minutes before each message.

import pandas as pd

# Load your sample data (replace with your actual data loading code)
data = {
    'date_time': ['10-05-2014 19:36:39', '10-05-2014 19:46:42', '10-05-2014 19:47:45',
                  '10-05-2014 19:48:48', '10-05-2014 19:50:14', '10-05-2014 19:54:44',
                  '10-05-2014 20:08:16'],
    'name': ['John', 'Pete', 'Joe', 'Mike', 'Aaron', 'Brad', 'Mike'],
    'text': ['Hi all', 'Hey', 'How are you', 'Good', 'Fine', 'What are you doing', 'Nothing']
}
df = pd.DataFrame(data)

# Convert date_time to datetime format
df['date_time'] = pd.to_datetime(df['date_time'], format='%d-%m-%Y %H:%M:%S')
# Sort data by time (critical for accurate window checks)
df = df.sort_values('date_time').reset_index(drop=True)

# Function to get participants in the 15-minute window before current message
def get_recent_participants(row):
    window_start = row['date_time'] - pd.Timedelta(minutes=15)
    # Filter messages in the window (excluding the current message itself)
    window_messages = df[(df['date_time'] >= window_start) & (df['date_time'] < row['date_time'])]
    # Get unique members, ordered by first appearance in the window
    unique_members = window_messages['name'].unique().tolist()
    return ', '.join(unique_members) if unique_members else pd.NA

# Apply function to create the in_convo column
df['in_convo'] = df.apply(get_recent_participants, axis=1)

This will produce exactly the in_convo column shown in your example.

Second: Individual Boolean Columns per Member (Best for Random Forest)

Since Random Forest models perform better with structured numerical features, creating a True/False column for each group member is the optimal choice. We can use pandas' rolling time window functionality to make this efficient, even with larger datasets.

# Create dummy variables where each column represents a member sending a message
member_dummies = pd.get_dummies(df['name'])

# Use a 15-minute rolling window to check if each member sent a message in the window
# `closed='left'` ensures we don't include the current message in the window
recent_participation = member_dummies.rolling('15min', on='date_time', closed='left').sum() > 0

# Rename columns for clarity and merge with original dataframe
recent_participation = recent_participation.add_prefix('has_')
df = pd.concat([df, recent_participation], axis=1)

This will add columns like has_John, has_Pete, etc., where each value is True if that member sent a message in the 15 minutes before the current message, and False otherwise. These columns are ready to feed directly into your Random Forest model without additional preprocessing.

Quick Notes:

  • Always sort your data by date_time before running these operations to ensure accurate window calculations.
  • The rolling window method is significantly faster than row-wise apply for large datasets, so prioritize this approach if you're working with more than just sample data.

内容的提问来源于stack exchange,提问作者DDO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 04:39:06