You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Coursera《Python数据科学导论》作业1正则表达式问题求助

Fixing Your Log Parsing Regex in Python

Let's break down what's wrong with your current regex pattern and fix it to correctly extract the required fields from your log data.

What's Wrong with Your Original Pattern

Your regex has several key issues that prevent it from matching the log lines:

  1. Host Group: You're trying to capture the IP address but your named host group uses \w* (which matches word characters, not IPs) while the actual IP pattern is outside the group. This means the host group captures nothing.
  2. User Name Group: Your pattern doesn't account for the - case (when no username is present), and the named group is set to \w* (capturing empty) instead of the actual username value.
  3. Time Group: The time string is inside square brackets, but your time group captures \w* (nothing) while the actual time is in an unnamed sub-group.
  4. Request Group: Similar to time, the request is inside quotes but your request group captures nothing, leaving the actual request in an unnamed group.

Corrected Code with Working Regex

Here's the fixed version of your function, with a regex that properly matches all parts of the log lines:

import re

def logs():
    with open("assets/logdata.txt", "r") as file:
        logdata = file.read()
    
    # Corrected regex pattern with named groups
    pattern = r"""
    ^(?P<host>\d+\.\d+\.\d+\.\d+)  # Capture the IP address as 'host'
    \s-\s                         # Match the "-" separator between host and user
    (?P<user_name>-|\w+)          # Capture username (either "-" or alphanumeric)
    \s\[                          # Match the opening bracket for time
    (?P<time>[^\]]+)              # Capture everything inside the brackets as 'time'
    \]\s"                         # Match closing bracket and opening quote
    (?P<request>[^"]+)            # Capture everything inside quotes as 'request'
    "
    """
    
    # Use MULTILINE to match each line, VERBOSE for readability
    matches = re.finditer(pattern, logdata, re.VERBOSE | re.MULTILINE)
    
    result = []
    for item in matches:
        parsed_dict = item.groupdict()
        print(parsed_dict)
        result.append(parsed_dict)
    
    return result

# Run the function to test
logs()

Key Improvements

  • Host: Directly matches IP addresses with \d+\.\d+\.\d+\.\d+ and captures it in the host group.
  • User Name: Handles both cases (- or a username) with (?P<user_name>-|\w+).
  • Time: Captures the full time string inside brackets using [^\]]+ (matches everything until the closing ]).
  • Request: Captures the full request inside quotes using [^"]+ (matches everything until the closing ").
  • Multiline Support: Added re.MULTILINE so ^ matches the start of each line in the log file.

Sample Output

For your test log lines, this will produce:

{"host":"159.253.153.40", "user_name":"-", "time":"21/Jun/2019:15:46:10 -0700", "request":"POST /e-business HTTP/1.0"}
{"host":"136.195.158.6", "user_name":"feeney9464", "time":"21/Jun/2019:15:46:11 -0700", "request":"HEAD /open-source/markets HTTP/2.0"}

内容的提问来源于stack exchange,提问作者BryantHsiung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 16:52:50