You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从日志文件提取邮箱地址?代码优化与问题排查请求

作业代码问题分析与指引

原代码与需求

原代码

'''This program is to read through any number of inputs (that is only: .txt files) the user passes through the sys.argv (through the terminal only). The file should only be a .txt file. Which means there is a conditional. Then the program should print those results to the terminal.'''
import sys

def find_email(line1):
    '''
    We must find the 'email address pattern' that is: username@domainname.domain
    There is only one symbol to help the read file method catch and print those email addresses:
    The @ symbol and then the two ' ' spaces at each end of the email address. 
    when we call from main() we must read the log file containing emails. 
    '''
    line1 = [' ', '@', ' ']
    at_postion = line1.find('@')
    first_place = line1.find(' ', at_postion)
    second_place = line1.find(' ', at_postion)
    return line1[second_place: first_place + 1]

def main():
    '''
    Main for sys argv input and one wrong file conditional.
    '''
    result = []
    results = []
    sys.argv(input('Enter .txt file: '))
    with open('r', sys.argv) as input:
        for result in find_email():
            result.append(results)
        print(result)
        if sys.argv(input('Enter .txt file: ')) == type(str):
            find_email()
        else:
            print("Error, enter .txt files only")
            exit()
    print(results)

if __name__ == "__main__":
    main()

需求说明

要求仅使用字符串方法从日志文件读取并提取垃圾邮箱地址至终端,具体需求:

  • a) 正确使用sys.argv,支持多文件输入,输入错误时显示提示;
  • b) 从指定日志文件提取所有邮箱,禁止使用静态文件;
  • c) 基于代码模块化重构代码。

代码理解误区与缺失点

  • sys.argv使用完全错误:sys.argv是命令行参数列表,不是可调用函数,不能用sys.argv(input(...))获取输入。正确方式是通过终端传参(如python script.py file1.txt file2.txt),然后取sys.argv[1:]获取所有传入的文件名,sys.argv[0]是脚本自身的名字。
  • find_email函数逻辑混乱:
    • 刚接收参数line1就把它覆盖成列表[' ', '@', ' '],完全丢弃了传入的文本行;
    • 列表没有find()方法,find()是字符串的方法,此处类型错误;
    • 提取邮箱的逻辑错误:first_place和second_place用相同的查找方式,无法定位邮箱的前后空格边界,切片结果无效。
  • 文件读取错误:
    • open()函数参数顺序颠倒,正确格式是open(filename, 'r'),不是open('r', filename);
    • 直接把sys.argv(列表)传给open(),应该遍历sys.argv[1:]中的每个文件名,逐个打开。
  • 列表操作错误:result.append(results)搞反了列表追加的方向,应该是results.append(result);且循环for result in find_email()没有传入参数,find_email()需要接收每行文本作为参数。
  • 输入验证逻辑无效:判断sys.argv(...) == type(str)毫无意义,应该检查文件名是否以.txt结尾,同时要处理文件不存在、无法读取的异常情况。

资源指引

1) 代码模块化

将功能拆分为独立、单一职责的函数:

  • 参数校验函数:接收文件名列表,过滤出非.txt文件,返回有效文件列表并提示无效项;
  • 邮箱提取函数:接收单文本行,用字符串方法定位邮箱并返回;
  • 文件处理函数:接收单个文件名,读取内容并调用邮箱提取函数,返回该文件的所有邮箱;
  • 主函数:处理sys.argv参数、调用上述函数、汇总结果并输出。

示例模块化结构:

import sys

def validate_files(filenames):
    valid = []
    for fn in filenames:
        if fn.endswith('.txt'):
            valid.append(fn)
        else:
            print(f"警告:{fn} 不是.txt文件,已跳过")
    return valid

def extract_email(line):
    at_pos = line.find('@')
    if at_pos == -1:
        return None
    # 向前找空格
    start = line.rfind(' ', 0, at_pos) + 1
    # 向后找空格
    end = line.find(' ', at_pos)
    if end == -1:
        end = len(line)
    return line[start:end].strip()

def process_file(filename):
    emails = []
    try:
        with open(filename, 'r', encoding='utf-8') as f:
            for line in f:
                email = extract_email(line)
                if email:
                    emails.append(email)
        return emails
    except FileNotFoundError:
        print(f"错误:文件 {filename} 不存在")
        return []
    except PermissionError:
        print(f"错误:无权限读取文件 {filename}")
        return []

def main():
    if len(sys.argv) < 2:
        print("使用方式:python script.py file1.txt [file2.txt ...]")
        return
    filenames = sys.argv[1:]
    valid_files = validate_files(filenames)
    all_emails = []
    for fn in valid_files:
        emails = process_file(fn)
        all_emails.extend(emails)
    print("提取到的邮箱:")
    for email in all_emails:
        print(email)

if __name__ == "__main__":
    main()

2) 文件读取方法

  • 逐行读取:用for line in file_object遍历文件,适合大文件,避免内存占用过高;
  • 上下文管理器:始终用with语句打开文件,它会自动关闭文件,避免资源泄漏;
  • 异常处理:捕获FileNotFoundError(文件不存在)、PermissionError(无权限)等常见异常,给出清晰提示;
  • 编码指定:打开文件时指定encoding='utf-8',避免编码解析错误。

3) 代码正确之处

  • 尝试将功能拆分为find_email和main函数,具备模块化的初步意识;
  • 知道使用with语句处理文件,能自动管理文件资源;
  • 有输入验证的意识,虽然逻辑错误,但核心需求的理解是对的。

4) 遗漏知识点

  • sys.argv的本质:它是一个列表,存储终端执行脚本时传入的所有参数,第一个元素是脚本路径;
  • 字符串方法细节:find()返回匹配的索引,找不到时返回-1,需要做判断;rfind()是从右往左查找,适合定位邮箱的起始空格;
  • 列表操作:append()是向列表末尾加单个元素,extend()是合并另一个列表的元素;
  • 异常处理:文件操作中必须处理异常,避免程序崩溃;
  • 切片逻辑:提取子串时要考虑边界情况(比如邮箱在行尾,找不到后空格的处理)。

内容的提问来源于stack exchange,提问作者JoelFU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:43:20