Coursera《Python数据科学导论》作业1正则表达式问题求助
Fixing Your Log Parsing Regex in Python
Let's break down what's wrong with your current regex pattern and fix it to correctly extract the required fields from your log data.
What's Wrong with Your Original Pattern
Your regex has several key issues that prevent it from matching the log lines:
- Host Group: You're trying to capture the IP address but your named
hostgroup uses\w*(which matches word characters, not IPs) while the actual IP pattern is outside the group. This means thehostgroup captures nothing. - User Name Group: Your pattern doesn't account for the
-case (when no username is present), and the named group is set to\w*(capturing empty) instead of the actual username value. - Time Group: The time string is inside square brackets, but your
timegroup captures\w*(nothing) while the actual time is in an unnamed sub-group. - Request Group: Similar to time, the request is inside quotes but your
requestgroup captures nothing, leaving the actual request in an unnamed group.
Corrected Code with Working Regex
Here's the fixed version of your function, with a regex that properly matches all parts of the log lines:
import re def logs(): with open("assets/logdata.txt", "r") as file: logdata = file.read() # Corrected regex pattern with named groups pattern = r""" ^(?P<host>\d+\.\d+\.\d+\.\d+) # Capture the IP address as 'host' \s-\s # Match the "-" separator between host and user (?P<user_name>-|\w+) # Capture username (either "-" or alphanumeric) \s\[ # Match the opening bracket for time (?P<time>[^\]]+) # Capture everything inside the brackets as 'time' \]\s" # Match closing bracket and opening quote (?P<request>[^"]+) # Capture everything inside quotes as 'request' " """ # Use MULTILINE to match each line, VERBOSE for readability matches = re.finditer(pattern, logdata, re.VERBOSE | re.MULTILINE) result = [] for item in matches: parsed_dict = item.groupdict() print(parsed_dict) result.append(parsed_dict) return result # Run the function to test logs()
Key Improvements
- Host: Directly matches IP addresses with
\d+\.\d+\.\d+\.\d+and captures it in thehostgroup. - User Name: Handles both cases (
-or a username) with(?P<user_name>-|\w+). - Time: Captures the full time string inside brackets using
[^\]]+(matches everything until the closing]). - Request: Captures the full request inside quotes using
[^"]+(matches everything until the closing"). - Multiline Support: Added
re.MULTILINEso^matches the start of each line in the log file.
Sample Output
For your test log lines, this will produce:
{"host":"159.253.153.40", "user_name":"-", "time":"21/Jun/2019:15:46:10 -0700", "request":"POST /e-business HTTP/1.0"} {"host":"136.195.158.6", "user_name":"feeney9464", "time":"21/Jun/2019:15:46:11 -0700", "request":"HEAD /open-source/markets HTTP/2.0"}
内容的提问来源于stack exchange,提问作者BryantHsiung
相关产品推荐
相关产品推荐

