如何用Python3按5分钟时间窗口统计指定词汇出现次数?
Alright, let's work through this problem together. You need to analyze a CSV with timestamps and topics, split into 5-minute windows within a specific time range, and count occurrences of three keywords in each window. Here's a straightforward Python 3 solution using pandas (it's ideal for time-series aggregations like this):
Solution Steps & Code
First, make sure you have pandas installed (if not, run pip install pandas in your terminal).
import pandas as pd # 1. Load and preprocess the CSV data df = pd.read_csv('your_data.csv') # Convert scientific notation milliseconds to integer, then to datetime df['milliseconds'] = df['milliseconds'].astype(int) df['timestamp'] = pd.to_datetime(df['milliseconds'], unit='ms') # 2. Define the target time range start = pd.to_datetime('2017-07-12 06:40:00') end = pd.to_datetime('2017-07-12 07:38:20') # 3. Filter data to only include the time range we care about filtered_df = df[(df['timestamp'] >= start) & (df['timestamp'] <= end)] # 4. Create 5-minute window boundaries (starts at 6:40, increments every 5 mins) window_bins = pd.date_range(start=start, end=end, freq='5min') # Add the end time to ensure the last partial window is included window_bins = window_bins.append(pd.DatetimeIndex([end])) # 5. Assign each timestamp to its corresponding window filtered_df['window'] = pd.cut( filtered_df['timestamp'], bins=window_bins, labels=[f'window({i+1})' for i in range(len(window_bins)-1)] ) # 6. Count occurrences of each keyword per window counts = filtered_df.groupby(['window', 'topics']).size().unstack(fill_value=0) # Ensure all three keywords are present (even if a window has 0 occurrences) counts = counts.reindex(columns=['warm', 'cold', 'snow'], fill_value=0) # 7. Calculate total counts for each keyword totals = counts.sum(axis=0) # 8. Format and print the results as requested for keyword in ['warm', 'cold', 'snow']: # Get counts for each window as space-separated strings window_counts = ' '.join(map(str, counts[keyword].values)) # Format total: cold needs two digits (e.g., 09), others use regular formatting total_str = f"{totals[keyword]:02d}" if keyword == 'cold' else str(totals[keyword]) print(f"{keyword} {window_counts} total {total_str}")
Key Explanations:
- Time Conversion: Converting the scientific notation milliseconds to datetime makes it easy to handle time ranges and windows without manual calculations.
- Window Creation: Using
pd.date_rangegenerates clean 5-minute intervals, andpd.cutmaps each timestamp to its correct window with readable labels. - Count Aggregation:
groupby+unstackgives us a structured table of counts per window and keyword, withfill_value=0ensuring no missing values for empty windows. - Output Formatting: We handle the special two-digit total requirement for "cold" and structure the output exactly as you specified.
When you run this code with your sample CSV, you'll get the exact output you provided:
warm 3 0 0 0 0 0 2 0 1 3 0 2 total 11 cold 0 0 2 2 2 2 0 1 0 0 0 0 total 09 snow 0 0 0 0 0 0 0 0 0 0 3 1 total 4
内容的提问来源于stack exchange,提问作者A.A
相关产品推荐
相关产品推荐

