You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

编写Python脚本读取日志文件找出404未找到的URL列表

Got it, here's a practical Python script to identify all URLs that returned a 404 status code from your Heroku-style log files. I’ve tailored this to work with the log structure you provided, including handling truncated lines like your second sample entry.

Python Script to Extract 404 URLs from Logs

Script Code

import re
from typing import List

def extract_404_urls(log_file_path: str) -> List[str]:
    """Extract all URLs (path values) that returned a 404 status code from the log file."""
    # Regex patterns to match status=404 and the path value
    status_pattern = re.compile(r'status=404')
    path_pattern = re.compile(r'path="([^"]+)"')
    urls_404 = []
    
    with open(log_file_path, 'r') as log_file:
        for line_num, line in enumerate(log_file, 1):
            line = line.strip()
            if not line:
                continue
            
            # Check if this line has a 404 status
            if status_pattern.search(line):
                # Extract the path URL
                path_match = path_pattern.search(line)
                if path_match:
                    # Extract host value to build full URL
                    try:
                        host = line.split('host=')[1].split()[0]
                        full_url = f"https://{host}{path_match.group(1)}"
                        urls_404.append(full_url)
                    except IndexError:
                        print(f"Warning: Line {line_num} has status=404 but no valid host found")
                else:
                    print(f"Warning: Line {line_num} has status=404 but no valid path found")
    
    return urls_404

if __name__ == "__main__":
    # Replace with your actual log file path
    LOG_FILE = "your_log_file.log"
    # Get all 404 URLs
    error_urls = extract_404_urls(LOG_FILE)
    
    # Print results to console
    print(f"Found {len(error_urls)} URLs returning 404:\n")
    for url in error_urls:
        print(f"- {url}")
    
    # Optional: Save results to a text file
    with open("404_urls.txt", 'w') as output_file:
        output_file.write("\n".join(error_urls))
    print(f"\nResults saved to 404_urls.txt")

Key Features

  • Robust Parsing: Uses regex to reliably match the status=404 flag and extract the URL path, even if log lines have varying whitespace or truncated content.
  • Full URL Construction: Combines the host value from the log with the path to create a complete, usable URL (e.g., https://workabledemo.com/api/accounts/3 from your first sample entry).
  • Error Handling: Warns about malformed lines that have a 404 status but missing host/path values.
  • Dual Output: Prints results to the console and optionally saves them to a 404_urls.txt file for later analysis.

How to Use

  1. Replace the LOG_FILE variable with the path to your actual log file (e.g., "heroku_router.log").
  2. Run the script with python extract_404s.py (or whatever you name the script file).
  3. Check the console output or the generated 404_urls.txt file for all 404 URLs.

Testing with Your Sample Logs

When run against your first sample entry, the script will output:

Found 1 URLs returning 404:

- https://workabledemo.com/api/accounts/3

Results saved to 404_urls.txt

It will also gracefully skip or warn about truncated lines like your second sample entry if it can't parse all required fields.

内容的提问来源于stack exchange,提问作者zee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:30:09