基于时间戳将TXT事件数据按24小时间隔分段的技术求助
Hey there! Let's finish that code to split your event data into 24-hour window segments. Here's a complete implementation that follows your requirement:
Step-by-Step Explanation
The core idea is to:
- Start with the first event's start time as the beginning of our first 24-hour window.
- Add all subsequent events to the current segment if their start time falls within this window.
- When we hit an event outside the window, we finalize the current segment and start a new one with this event's start time as the new window's beginning.
Complete Python Code
import datetime from datetime import timedelta import locale def split_events_by_24h_window(file_path): # Set locale to English to parse month abbreviations like "MAR" correctly try: locale.setlocale(locale.LC_TIME, 'en_US.UTF-8') # Unix/macOS except: try: locale.setlocale(locale.LC_TIME, 'English_US') # Windows except: print("Warning: Could not set English locale. Month parsing may fail.") segments = [] with open(file_path, 'r') as input_file: lines = [line.strip() for line in input_file.readlines() if line.strip()] if not lines: return segments # Initialize first segment with the first event first_parts = lines[0].split() start_time_str = f"{first_parts[2]} {first_parts[3]}" current_window_start = datetime.datetime.strptime(start_time_str, '%d%b%y %H%M%S') current_window_end = current_window_start + timedelta(hours=24) current_segment = [lines[0]] # Process remaining events for line in lines[1:]: parts = line.split() event_start_str = f"{parts[2]} {parts[3]}" event_start = datetime.datetime.strptime(event_start_str, '%d%b%y %H%M%S') if event_start <= current_window_end: current_segment.append(line) else: # Finalize current segment and start a new one segments.append(current_segment) current_segment = [line] current_window_start = event_start current_window_end = current_window_start + timedelta(hours=24) # Add the last segment to the list segments.append(current_segment) return segments # Example usage if __name__ == "__main__": event_segments = split_events_by_24h_window('file.txt') # Print each segment for verification for idx, segment in enumerate(event_segments, 1): print(f"Segment {idx}:") for event in segment: print(f" {event}") print("-" * 40)
Key Details
- Locale Handling: The code tries to set an English locale to ensure month abbreviations like "MAR" are parsed correctly (critical for cross-platform compatibility).
- Empty Line Handling: Skips any empty lines in the input file to avoid errors.
- Segment Initialization: Starts with the first event's start time and creates a 24-hour window from that point.
- Window Check: For each subsequent event, checks if its start time is within the current window. If not, starts a new segment.
Testing with Your Data
When you run this code with your input file, it will produce segments like:
- Segment 1: Only the first event (15MAR18 103000)
- Segment 2: Events from 16MAR18 120518 up to 16MAR18 151733
- Segment 3: Events starting from 17MAR18 165013 up to 17MAR18 203738 (since this is within 24 hours of 17MAR18 165013)
- And so on for the rest of your events.
内容的提问来源于stack exchange,提问作者arsenal88
相关产品推荐
相关产品推荐

