如何处理CSV/Excel序列文件生成反向互补序列并修复Python脚本报错
错误根因
- 变量作用域问题:
inputfile、outputfile都是定义在main函数内部的局部变量,函数外部的文件读写代码无权访问这两个变量,因此触发未定义报错。 - 执行顺序错误:Python脚本会从上到下逐行执行全局作用域的代码,当前把文件读写逻辑写在了
main函数外,会在调用main函数解析参数之前就执行,此时两个文件路径还没完成赋值,自然无法使用。
修复方案
把文件读写逻辑全部移入main函数内部,同时优化两个细节:
- 读取每行序列后先去除末尾的换行符、空白字符再做反向互补计算,避免结果出现异常字符
- 如果需要输出标准Excel格式,可借助
pandas库实现,比直接写csv后改后缀更规范
基础修复版本(输出csv格式)
import sys, getopt def main(argv): inputfile = '' outputfile = '' try: opts, args = getopt.getopt(argv,"hi:o:",["ifile=","ofile="]) except getopt.GetoptError: print('reversecomplement2.py -i <input csv> -o <output csv>') sys.exit(2) for opt, arg in opts: if opt == '-h': print('reversecomplement2.py -i <input csv> -o <output csv>') sys.exit() elif opt in ("-i", "--ifile"): inputfile = arg elif opt in ("-o", "--ofile"): outputfile = arg # 把文件读写逻辑移到main函数内部 complement = {'A': 'T', 'C': 'G', 'G': 'C', 'T': 'A'} with open(outputfile, 'w') as final: with open(inputfile, 'r') as input: for line in input: # 先去掉换行符和首尾空白再计算 seq = line.strip() if not seq: continue reverse_complement = "".join(complement.get(base, base) for base in reversed(seq)) final.write(reverse_complement + '\n') if __name__ == "__main__": main(sys.argv[1:])
优化版本(直接输出Excel格式)
使用前先执行命令安装依赖:pip install pandas openpyxl
import sys, getopt import pandas as pd def main(argv): inputfile = '' outputfile = '' try: opts, args = getopt.getopt(argv,"hi:o:",["ifile=","ofile="]) except getopt.GetoptError: print('reversecomplement2.py -i <input csv> -o <output xlsx>') sys.exit(2) for opt, arg in opts: if opt == '-h': print('reversecomplement2.py -i <input csv> -o <output xlsx>') sys.exit() elif opt in ("-i", "--ifile"): inputfile = arg elif opt in ("-o", "--ofile"): outputfile = arg complement = {'A': 'T', 'C': 'G', 'G': 'C', 'T': 'A'} result = [] with open(inputfile, 'r') as input: for line in input: seq = line.strip() if not seq: continue reverse_complement = "".join(complement.get(base, base) for base in reversed(seq)) # 同时保留原始序列和反向互补序列,方便对照 result.append({"原始序列": seq, "反向互补序列": reverse_complement}) # 写入Excel文件 df = pd.DataFrame(result) df.to_excel(outputfile, index=False) if __name__ == "__main__": main(sys.argv[1:])
内容的提问来源于stack exchange,提问作者pythonbeginner
相关产品推荐
相关产品推荐

