如何在多个Python脚本中查找忽略空格与空行的重复代码行
查找Python脚本中重复的代码行(忽略首尾空格与空行)
针对你要在分散的Python脚本里找出重复代码行(比如字典键值对)、忽略首尾空格和空行的需求,我整理了两种实用方案,你可以根据自己的习惯选择:
方案一:Unix终端命令链
如果你熟悉Unix/Linux终端,用组合命令就能快速实现需求,不需要写额外脚本:
步骤1:提取所有有效行并记录位置
这条命令会递归遍历当前目录及子目录下的所有.py文件,过滤掉空行(包括全是空格的行),去掉每行首尾空格,同时记录每行的原始文件和行号:
grep -r -n '.' --include='*.py' . | sed '/^[[:space:]]*$/d' | awk '{gsub(/^[[:space:]]+|[[:space:]]+$/, "", $0); print $0 "\t" FILENAME ":" FNR}'
步骤2:筛选重复行
把上面的输出管道到sort和uniq,就能找出所有重复出现的行,以及它们的位置:
grep -r -n '.' --include='*.py' . | sed '/^[[:space:]]*$/d' | awk '{gsub(/^[[:space:]]+|[[:space:]]+$/, "", $0); print $0 "\t" FILENAME ":" FNR}' | sort | uniq -d --all-repeated=separate
sort:将处理后的行按内容排序,让重复行相邻uniq -d --all-repeated=separate:只显示重复出现的行,并且把每组重复行用空行分隔开,方便查看
方案二:自定义Python脚本
如果需要更灵活的过滤逻辑(比如只针对字典键值对行),可以写一个Python脚本,精准控制筛选规则:
import os from collections import defaultdict def find_duplicate_lines(root_dir): # 用字典存储:清理后的行 -> [文件位置列表] line_map = defaultdict(list) # 递归遍历所有.py文件 for dirpath, _, filenames in os.walk(root_dir): for filename in filenames: if filename.endswith('.py'): file_path = os.path.join(dirpath, filename) with open(file_path, 'r', encoding='utf-8') as f: # 遍历每一行,记录行号(从1开始) for line_num, line in enumerate(f, 1): cleaned_line = line.strip() if not cleaned_line: # 跳过空行 continue # 可选:添加过滤规则,比如只找字典键值对行 # if ':' in cleaned_line and (cleaned_line.startswith(("'", '"')) or cleaned_line.endswith(',')): line_map[cleaned_line].append(f"{file_path}:{line_num}") # 输出所有重复的行 print("找到以下重复代码行:") print("="*50) for line, locations in line_map.items(): if len(locations) > 1: print(f"重复内容: `{line}`") print("出现位置:") for loc in locations: print(f" - {loc}") print("-"*50) if __name__ == "__main__": # 传入要搜索的根目录,这里用当前目录 find_duplicate_lines('.')
脚本说明
- 自动跳过空行和全空格行,清理每行首尾空格后再判断重复
- 可以根据需求添加过滤规则(比如注释掉的那行代码,只筛选字典键值对格式的行)
- 输出清晰,会列出重复内容和所有出现的文件位置
后续合并重复字典的建议
找到重复的字典行后,你可以把同用途的字典合并到一个单独的模块(比如translations.py):
# translations.py GREETINGS = { 'de': 'hallo', 'en': 'hello', 'fa': 'سلام', }
然后在其他脚本里导入使用:
# hi.py from translations import GREETINGS hi_dict = GREETINGS # maintenance/hello.py from translations import GREETINGS class Hello(Superclass): hello_dict = GREETINGS
这样就能统一维护字典内容,降低代码冗余,也能提升CodeClimate评分。
内容的提问来源于stack exchange,提问作者aleskva
相关产品推荐
相关产品推荐

