如何用Python从日志文件提取修改密码达10次的用户信息
Solution to Track Users with 10+ Password Modification Attempts
Hey there! Let's work through your goal of extracting details for users who've had 10+ failed authentication attempts (which you've linked to password modification attempts). First, let's go over a couple of gaps in your current code, then share an updated version that meets your needs.
Issues in the Current Code
- Overwriting CSV: Using
open(output_file, 'w')inside the loop means you're overwriting the CSV file every time a matching line is found—only the last record will be saved. - No Count Tracking: There's no logic to count how many times each user has triggered the
MBX_AUTHENTICATION_FAILEDevent, so we can't filter for users with 10+ attempts. - Redundant Regex: You're running the same date/time regex twice; we can optimize that to run once per line.
Updated Code
Here's a revised version that fixes these issues and adds the count-tracking logic:
import re from csv import writer import datetime log_file = '/Users/kiya/Desktop/ip.txt' output_file = '/Users/kiya/Desktop/output.csv' name_to_check = 'MBX_AUTHENTICATION_FAILED' # Dictionary to track each user's failed attempts: key = username, value = list of (date, time, ip) user_records = {} with open(log_file, encoding="utf-8") as infile: for line in infile: if name_to_check not in line: continue # Skip lines that don't match our target # Extract username username_match = re.search(r'(?<=userName=\[)(.*?)(?=\],)', line) if not username_match: continue # Skip lines where username can't be extracted username = username_match.group() # Extract and format timestamp datetime_match = re.search(r'(?P<date>\d{8})\s+(?P<time>\d{9})\+(?P<zone>\d{4})', line) if not datetime_match: continue # Skip lines with invalid timestamp date_str = datetime.datetime.strptime(datetime_match.group('date'), "%Y%m%d").strftime("%Y-%m-%d") time_str = datetime.datetime.strptime(datetime_match.group('time'), "%H%M%S%f").strftime("%H:%M:%S") # Extract IP address ip_match = re.search(r'(([0-9]|[1-9][0-9]|1[0-9]{2}|2[0-4][0-9]|25[0-5])\.){3}([0-9]|[1-9][0-9]|1[0-9]{2}|2[0-4][0-9]|25[0-5])', line) if not ip_match: continue # Skip lines where IP can't be extracted ip = ip_match.group() # Add record to user's list if username not in user_records: user_records[username] = [] user_records[username].append((date_str, time_str, ip)) # Now write only users with 10+ attempts to CSV with open(output_file, 'w', newline='', encoding="utf-8") as outfile: csv_writer = writer(outfile) # Write header csv_writer.writerow(["Username", "Date", "Time", "IP"]) # Iterate through users and their records for username, records in user_records.items(): if len(records) >= 10: # Write each record for this user for date, time, ip in records: csv_writer.writerow([username, date, time, ip])
Key Improvements Explained
- Count Tracking: The
user_recordsdictionary keeps track of every failed attempt for each user, so we can easily check if they've hit 10+ attempts. - Non-Destructive CSV Writing: We open the output file once at the end, writing all qualifying records in one go (no overwriting).
- Error Handling: Added checks to skip lines where any field (username, timestamp, IP) can't be extracted—this prevents crashes from malformed log lines.
- Optimized Regex: The date/time regex runs once per line instead of twice, saving redundant processing.
How to Use
- Double-check that your log file path and output CSV path are correct.
- Run the script—it will process all log lines, count attempts per user, and export only users with 10+ attempts to the CSV file.
内容的提问来源于stack exchange,提问作者Katarina Alves
相关产品推荐
相关产品推荐

