基于pandas实现WhatsApp群聊时间窗口参与者提取,用于Random Forest回复者预测
Got it, let's tackle this problem step by step. You're working on predicting message responders in an 11-person WhatsApp group using Random Forest, and need to add features tracking which members participated in the conversation within a specified time window (15 minutes in your example). Below are implementations for both of your desired feature formats.
First: Single Column with Participant List
First off, make sure your date_time column is properly converted to a datetime type—this is critical for accurate time window calculations. Then we'll use a row-wise check to pull all unique members who sent messages in the 15 minutes before each message.
import pandas as pd # Load your sample data (replace with your actual data loading code) data = { 'date_time': ['10-05-2014 19:36:39', '10-05-2014 19:46:42', '10-05-2014 19:47:45', '10-05-2014 19:48:48', '10-05-2014 19:50:14', '10-05-2014 19:54:44', '10-05-2014 20:08:16'], 'name': ['John', 'Pete', 'Joe', 'Mike', 'Aaron', 'Brad', 'Mike'], 'text': ['Hi all', 'Hey', 'How are you', 'Good', 'Fine', 'What are you doing', 'Nothing'] } df = pd.DataFrame(data) # Convert date_time to datetime format df['date_time'] = pd.to_datetime(df['date_time'], format='%d-%m-%Y %H:%M:%S') # Sort data by time (critical for accurate window checks) df = df.sort_values('date_time').reset_index(drop=True) # Function to get participants in the 15-minute window before current message def get_recent_participants(row): window_start = row['date_time'] - pd.Timedelta(minutes=15) # Filter messages in the window (excluding the current message itself) window_messages = df[(df['date_time'] >= window_start) & (df['date_time'] < row['date_time'])] # Get unique members, ordered by first appearance in the window unique_members = window_messages['name'].unique().tolist() return ', '.join(unique_members) if unique_members else pd.NA # Apply function to create the in_convo column df['in_convo'] = df.apply(get_recent_participants, axis=1)
This will produce exactly the in_convo column shown in your example.
Second: Individual Boolean Columns per Member (Best for Random Forest)
Since Random Forest models perform better with structured numerical features, creating a True/False column for each group member is the optimal choice. We can use pandas' rolling time window functionality to make this efficient, even with larger datasets.
# Create dummy variables where each column represents a member sending a message member_dummies = pd.get_dummies(df['name']) # Use a 15-minute rolling window to check if each member sent a message in the window # `closed='left'` ensures we don't include the current message in the window recent_participation = member_dummies.rolling('15min', on='date_time', closed='left').sum() > 0 # Rename columns for clarity and merge with original dataframe recent_participation = recent_participation.add_prefix('has_') df = pd.concat([df, recent_participation], axis=1)
This will add columns like has_John, has_Pete, etc., where each value is True if that member sent a message in the 15 minutes before the current message, and False otherwise. These columns are ready to feed directly into your Random Forest model without additional preprocessing.
Quick Notes:
- Always sort your data by
date_timebefore running these operations to ensure accurate window calculations. - The rolling window method is significantly faster than row-wise
applyfor large datasets, so prioritize this approach if you're working with more than just sample data.
内容的提问来源于stack exchange,提问作者DDO

