You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

编写Python函数检测日志文件中指定文件名的重复出现情况

Hey there, let's build a simple yet flexible Python function to solve your problem. You want to check if specific strings (like first.arxml or second.arxml) show up multiple times in your log file—here's how to do it properly:

The Python Function

This function handles common edge cases like file not found errors, case sensitivity, and avoids partial matches (so you don't accidentally count first.arxml.bak as a match for first.arxml):

def check_string_duplicate(log_file_path, target_string, case_sensitive=True, whole_word=False):
    """
    Check if a target string appears multiple times in a log file
    
    Args:
        log_file_path (str): Path to your .log file
        target_string (str): The string you want to check (e.g., "first.arxml")
        case_sensitive (bool): Whether to respect uppercase/lowercase, default True
        whole_word (bool): Match only complete words (avoids partial hits), default False
    
    Returns:
        tuple: (total_occurrences, is_duplicated)
    """
    count = 0
    try:
        # Open the log file with UTF-8 encoding (adjust if your log uses a different one)
        with open(log_file_path, 'r', encoding='utf-8') as log_file:
            for line in log_file:
                # Handle case sensitivity if needed
                processed_line = line if case_sensitive else line.lower()
                processed_target = target_string if case_sensitive else target_string.lower()
                
                if whole_word:
                    # Use regex to match whole words only (escape special chars in target)
                    import re
                    matches = re.findall(rf'\b{re.escape(processed_target)}\b', processed_line)
                    count += len(matches)
                else:
                    # Simple substring count
                    count += processed_line.count(processed_target)
        
        # Determine if the string appeared more than once
        is_duplicated = count > 1
        return count, is_duplicated

    except FileNotFoundError:
        print(f"Error: Could not find log file at {log_file_path}")
        return 0, False
    except Exception as e:
        print(f"Error reading log file: {str(e)}")
        return 0, False

How to Use It

Here's a quick example to test with your files:

# Replace with your actual log file path
log_file = "your_generated.log"

# Check for first.arxml (match whole words to avoid partial hits)
occurrences, is_duplicated = check_string_duplicate(log_file, "first.arxml", whole_word=True)

if is_duplicated:
    print(f"'first.arxml' appears {occurrences} times in the log (duplicate found!)")
else:
    print(f"'first.arxml' appears {occurrences} time(s) in the log")

# Repeat for second.arxml
occurrences, is_duplicated = check_string_duplicate(log_file, "second.arxml", whole_word=True)

if is_duplicated:
    print(f"'second.arxml' appears {occurrences} times in the log (duplicate found!)")
else:
    print(f"'second.arxml' appears {occurrences} time(s) in the log")

Key Notes

  • Whole Word Matching: Set whole_word=True if you want to ensure you're only counting exact filename matches (this prevents false positives like first.arxml.backup being counted).
  • Case Sensitivity: If your log might have mixed case (e.g., First.arxml), set case_sensitive=False to ignore case differences.
  • Encoding: If your log file uses a non-UTF-8 encoding (like latin-1), adjust the encoding parameter in open() accordingly.

内容的提问来源于stack exchange,提问作者Robert

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:21:02