编写Python脚本读取日志文件找出404未找到的URL列表
Got it, here's a practical Python script to identify all URLs that returned a 404 status code from your Heroku-style log files. I’ve tailored this to work with the log structure you provided, including handling truncated lines like your second sample entry.
Python Script to Extract 404 URLs from Logs
Script Code
import re from typing import List def extract_404_urls(log_file_path: str) -> List[str]: """Extract all URLs (path values) that returned a 404 status code from the log file.""" # Regex patterns to match status=404 and the path value status_pattern = re.compile(r'status=404') path_pattern = re.compile(r'path="([^"]+)"') urls_404 = [] with open(log_file_path, 'r') as log_file: for line_num, line in enumerate(log_file, 1): line = line.strip() if not line: continue # Check if this line has a 404 status if status_pattern.search(line): # Extract the path URL path_match = path_pattern.search(line) if path_match: # Extract host value to build full URL try: host = line.split('host=')[1].split()[0] full_url = f"https://{host}{path_match.group(1)}" urls_404.append(full_url) except IndexError: print(f"Warning: Line {line_num} has status=404 but no valid host found") else: print(f"Warning: Line {line_num} has status=404 but no valid path found") return urls_404 if __name__ == "__main__": # Replace with your actual log file path LOG_FILE = "your_log_file.log" # Get all 404 URLs error_urls = extract_404_urls(LOG_FILE) # Print results to console print(f"Found {len(error_urls)} URLs returning 404:\n") for url in error_urls: print(f"- {url}") # Optional: Save results to a text file with open("404_urls.txt", 'w') as output_file: output_file.write("\n".join(error_urls)) print(f"\nResults saved to 404_urls.txt")
Key Features
- Robust Parsing: Uses regex to reliably match the
status=404flag and extract the URL path, even if log lines have varying whitespace or truncated content. - Full URL Construction: Combines the
hostvalue from the log with the path to create a complete, usable URL (e.g.,https://workabledemo.com/api/accounts/3from your first sample entry). - Error Handling: Warns about malformed lines that have a 404 status but missing host/path values.
- Dual Output: Prints results to the console and optionally saves them to a
404_urls.txtfile for later analysis.
How to Use
- Replace the
LOG_FILEvariable with the path to your actual log file (e.g.,"heroku_router.log"). - Run the script with
python extract_404s.py(or whatever you name the script file). - Check the console output or the generated
404_urls.txtfile for all 404 URLs.
Testing with Your Sample Logs
When run against your first sample entry, the script will output:
Found 1 URLs returning 404: - https://workabledemo.com/api/accounts/3 Results saved to 404_urls.txt
It will also gracefully skip or warn about truncated lines like your second sample entry if it can't parse all required fields.
内容的提问来源于stack exchange,提问作者zee
相关产品推荐
相关产品推荐

