编写Python函数检测日志文件中指定文件名的重复出现情况
Hey there, let's build a simple yet flexible Python function to solve your problem. You want to check if specific strings (like first.arxml or second.arxml) show up multiple times in your log file—here's how to do it properly:
The Python Function
This function handles common edge cases like file not found errors, case sensitivity, and avoids partial matches (so you don't accidentally count first.arxml.bak as a match for first.arxml):
def check_string_duplicate(log_file_path, target_string, case_sensitive=True, whole_word=False): """ Check if a target string appears multiple times in a log file Args: log_file_path (str): Path to your .log file target_string (str): The string you want to check (e.g., "first.arxml") case_sensitive (bool): Whether to respect uppercase/lowercase, default True whole_word (bool): Match only complete words (avoids partial hits), default False Returns: tuple: (total_occurrences, is_duplicated) """ count = 0 try: # Open the log file with UTF-8 encoding (adjust if your log uses a different one) with open(log_file_path, 'r', encoding='utf-8') as log_file: for line in log_file: # Handle case sensitivity if needed processed_line = line if case_sensitive else line.lower() processed_target = target_string if case_sensitive else target_string.lower() if whole_word: # Use regex to match whole words only (escape special chars in target) import re matches = re.findall(rf'\b{re.escape(processed_target)}\b', processed_line) count += len(matches) else: # Simple substring count count += processed_line.count(processed_target) # Determine if the string appeared more than once is_duplicated = count > 1 return count, is_duplicated except FileNotFoundError: print(f"Error: Could not find log file at {log_file_path}") return 0, False except Exception as e: print(f"Error reading log file: {str(e)}") return 0, False
How to Use It
Here's a quick example to test with your files:
# Replace with your actual log file path log_file = "your_generated.log" # Check for first.arxml (match whole words to avoid partial hits) occurrences, is_duplicated = check_string_duplicate(log_file, "first.arxml", whole_word=True) if is_duplicated: print(f"'first.arxml' appears {occurrences} times in the log (duplicate found!)") else: print(f"'first.arxml' appears {occurrences} time(s) in the log") # Repeat for second.arxml occurrences, is_duplicated = check_string_duplicate(log_file, "second.arxml", whole_word=True) if is_duplicated: print(f"'second.arxml' appears {occurrences} times in the log (duplicate found!)") else: print(f"'second.arxml' appears {occurrences} time(s) in the log")
Key Notes
- Whole Word Matching: Set
whole_word=Trueif you want to ensure you're only counting exact filename matches (this prevents false positives likefirst.arxml.backupbeing counted). - Case Sensitivity: If your log might have mixed case (e.g.,
First.arxml), setcase_sensitive=Falseto ignore case differences. - Encoding: If your log file uses a non-UTF-8 encoding (like
latin-1), adjust theencodingparameter inopen()accordingly.
内容的提问来源于stack exchange,提问作者Robert
相关产品推荐
相关产品推荐

