如何用Python删除文件中每个OPEN(后首个前置逗号(限定在)CLOSE前)
解决:删除每个
OPEN(后对应) CLOSE内的第一个逗号 需求
处理文本文件,把每个OPEN(开头、) CLOSE结尾的块里,第一个出现的逗号删掉。
示例
输入:
OPEN( ,a ,b ) CLOSE OPEN( c ) CLOSE OPEN( ,d ,e ,f ) CLOSE
处理后输出:
OPEN( a ,b ) CLOSE OPEN( c ) CLOSE OPEN( d ,e ,f ) CLOSE
我试过的方法
一开始想写正则,但不知道怎么精准匹配这种跨行的块,还考虑过用awk/sed。试了个不对的正则(当时混淆了别的需求):
pattern = r'WITH \(([^)]+),([^)]+)\) AS' replacement = r'WITH (\1\2) AS' sql_content_modified = re.sub(pattern, replacement, sql_content)
最后用了字符串替换,但感觉太局限,只能处理逗号在(下一行的情况:
# 读取SQL文件 with open(f'{filename}', 'r') as file: content = file.read() content_modified = content.replace('(\n,', '(') content_modified = re.sub('--<([a-z]*.*[A-Z]*)>', '', content) # 删除--<*>行 # 写回文件 with open(f'{filename}', 'w') as file: file.write(content_modified) remove_empty_lines_from_file(filename) # 再处理一次WITH开头的情况 with open(f'{filename}', 'r') as file: content = file.read() content_modified = content.replace('(\n,', '(\n') with open(f'{filename}', 'w') as file: file.write(content_modified)
靠谱的解决办法
1. Python正则(小文件首选)
用Python正则的DOTALL模式匹配整个OPEN(...) CLOSE块,然后在每个块里删掉第一个逗号:
import re def fix_file(filename): with open(filename, 'r') as f: content = f.read() # 处理每个匹配到的OPEN块 def handle_block(match): inner_content = match.group(1) # 只删除第一个逗号 cleaned_inner = inner_content.replace(',', '', 1) return f'OPEN({cleaned_inner}) CLOSE' # 跨行匹配所有OPEN(...) CLOSE块 modified_content = re.sub(r'OPEN\((.*?)\) CLOSE', handle_block, content, flags=re.DOTALL) with open(filename, 'w') as f: f.write(modified_content)
2. 状态机逐行处理(大文件适用)
如果文件太大,没法一次性读进内存,就用状态跟踪逐行处理:
def fix_large_file(filename): in_open_block = False has_removed_comma = False output = [] with open(filename, 'r') as f: for line in f: line_stripped = line.strip() if line_stripped.startswith('OPEN('): in_open_block = True has_removed_comma = False output.append(line) elif line_stripped == ') CLOSE': in_open_block = False output.append(line) elif in_open_block and not has_removed_comma: # 删掉这行里的第一个逗号 fixed_line = line.replace(',', '', 1) output.append(fixed_line) has_removed_comma = True else: output.append(line) with open(filename, 'w') as f: f.writelines(output)
3. awk命令行处理
不想写Python脚本的话,用awk一行搞定:
awk ' /OPEN\(/ { in_block=1; comma_removed=0; print; next } /\) CLOSE/ { in_block=0; print; next } in_block && !comma_removed { sub(/,/, "", $0); comma_removed=1 } 1' input.txt > output.txt
说明
- 正则方案代码简洁,适合小文件;
- 状态机方案内存占用低,适合大文件;
- awk方案适合命令行快速批量处理。
内容的提问来源于stack exchange,提问作者SquatLicense
相关产品推荐
相关产品推荐

