如何从日志文件提取邮箱地址?代码优化与问题排查请求
作业代码问题分析与指引
原代码与需求
原代码
'''This program is to read through any number of inputs (that is only: .txt files) the user passes through the sys.argv (through the terminal only). The file should only be a .txt file. Which means there is a conditional. Then the program should print those results to the terminal.''' import sys def find_email(line1): ''' We must find the 'email address pattern' that is: username@domainname.domain There is only one symbol to help the read file method catch and print those email addresses: The @ symbol and then the two ' ' spaces at each end of the email address. when we call from main() we must read the log file containing emails. ''' line1 = [' ', '@', ' '] at_postion = line1.find('@') first_place = line1.find(' ', at_postion) second_place = line1.find(' ', at_postion) return line1[second_place: first_place + 1] def main(): ''' Main for sys argv input and one wrong file conditional. ''' result = [] results = [] sys.argv(input('Enter .txt file: ')) with open('r', sys.argv) as input: for result in find_email(): result.append(results) print(result) if sys.argv(input('Enter .txt file: ')) == type(str): find_email() else: print("Error, enter .txt files only") exit() print(results) if __name__ == "__main__": main()
需求说明
要求仅使用字符串方法从日志文件读取并提取垃圾邮箱地址至终端,具体需求:
- a) 正确使用
sys.argv,支持多文件输入,输入错误时显示提示; - b) 从指定日志文件提取所有邮箱,禁止使用静态文件;
- c) 基于代码模块化重构代码。
代码理解误区与缺失点
sys.argv使用完全错误:sys.argv是命令行参数列表,不是可调用函数,不能用sys.argv(input(...))获取输入。正确方式是通过终端传参(如python script.py file1.txt file2.txt),然后取sys.argv[1:]获取所有传入的文件名,sys.argv[0]是脚本自身的名字。find_email函数逻辑混乱:- 刚接收参数
line1就把它覆盖成列表[' ', '@', ' '],完全丢弃了传入的文本行; - 列表没有
find()方法,find()是字符串的方法,此处类型错误; - 提取邮箱的逻辑错误:
first_place和second_place用相同的查找方式,无法定位邮箱的前后空格边界,切片结果无效。
- 刚接收参数
- 文件读取错误:
open()函数参数顺序颠倒,正确格式是open(filename, 'r'),不是open('r', filename);- 直接把
sys.argv(列表)传给open(),应该遍历sys.argv[1:]中的每个文件名,逐个打开。
- 列表操作错误:
result.append(results)搞反了列表追加的方向,应该是results.append(result);且循环for result in find_email()没有传入参数,find_email()需要接收每行文本作为参数。 - 输入验证逻辑无效:判断
sys.argv(...) == type(str)毫无意义,应该检查文件名是否以.txt结尾,同时要处理文件不存在、无法读取的异常情况。
资源指引
1) 代码模块化
将功能拆分为独立、单一职责的函数:
- 参数校验函数:接收文件名列表,过滤出非
.txt文件,返回有效文件列表并提示无效项; - 邮箱提取函数:接收单文本行,用字符串方法定位邮箱并返回;
- 文件处理函数:接收单个文件名,读取内容并调用邮箱提取函数,返回该文件的所有邮箱;
- 主函数:处理
sys.argv参数、调用上述函数、汇总结果并输出。
示例模块化结构:
import sys def validate_files(filenames): valid = [] for fn in filenames: if fn.endswith('.txt'): valid.append(fn) else: print(f"警告:{fn} 不是.txt文件,已跳过") return valid def extract_email(line): at_pos = line.find('@') if at_pos == -1: return None # 向前找空格 start = line.rfind(' ', 0, at_pos) + 1 # 向后找空格 end = line.find(' ', at_pos) if end == -1: end = len(line) return line[start:end].strip() def process_file(filename): emails = [] try: with open(filename, 'r', encoding='utf-8') as f: for line in f: email = extract_email(line) if email: emails.append(email) return emails except FileNotFoundError: print(f"错误:文件 {filename} 不存在") return [] except PermissionError: print(f"错误:无权限读取文件 {filename}") return [] def main(): if len(sys.argv) < 2: print("使用方式:python script.py file1.txt [file2.txt ...]") return filenames = sys.argv[1:] valid_files = validate_files(filenames) all_emails = [] for fn in valid_files: emails = process_file(fn) all_emails.extend(emails) print("提取到的邮箱:") for email in all_emails: print(email) if __name__ == "__main__": main()
2) 文件读取方法
- 逐行读取:用
for line in file_object遍历文件,适合大文件,避免内存占用过高; - 上下文管理器:始终用
with语句打开文件,它会自动关闭文件,避免资源泄漏; - 异常处理:捕获
FileNotFoundError(文件不存在)、PermissionError(无权限)等常见异常,给出清晰提示; - 编码指定:打开文件时指定
encoding='utf-8',避免编码解析错误。
3) 代码正确之处
- 尝试将功能拆分为
find_email和main函数,具备模块化的初步意识; - 知道使用
with语句处理文件,能自动管理文件资源; - 有输入验证的意识,虽然逻辑错误,但核心需求的理解是对的。
4) 遗漏知识点
sys.argv的本质:它是一个列表,存储终端执行脚本时传入的所有参数,第一个元素是脚本路径;- 字符串方法细节:
find()返回匹配的索引,找不到时返回-1,需要做判断;rfind()是从右往左查找,适合定位邮箱的起始空格; - 列表操作:
append()是向列表末尾加单个元素,extend()是合并另一个列表的元素; - 异常处理:文件操作中必须处理异常,避免程序崩溃;
- 切片逻辑:提取子串时要考虑边界情况(比如邮箱在行尾,找不到后空格的处理)。
内容的提问来源于stack exchange,提问作者JoelFU
相关产品推荐
相关产品推荐

