后端报错‘dict’对象无‘lower’属性,求钓鱼邮件检测模型解决方案
问题原因
报错'dict' object has no attribute 'lower'的核心是**TfidfVectorizer在处理文本时收到了字典对象,而非预期的字符串**。
从报错栈可以看到,sklearn的文本预处理环节尝试调用doc.lower()转换小写,但传入的doc是字典,字典没有lower方法,因此抛出异常。
修复方案
1. 修正输入数据的获取与处理逻辑
你的后端代码中,email_content = payload.get('email_content', None)可能没有拿到正确的字符串:要么前端传入的email_content字段值是字典,要么payload结构不符合预期。
修改后端预测接口代码
在使用email_content前,增加类型检查和处理:
@app.route('/predict', methods=['POST']) def predict_phishing(): if request.method == 'POST': try: payload = request.json email_content = payload.get('email_content', None) # 检查必填字段是否存在 if email_content is None: return jsonify({'error': '缺少email_content字段'}), HTTPStatus.BAD_REQUEST # 处理字典类型的email_content(比如前端拆分了主题和正文) if isinstance(email_content, dict): # 根据实际字段拼接,示例假设包含subject和body email_content = f"{email_content.get('subject', '')} {email_content.get('body', '')}" # 强制转换为字符串,避免非字符串类型传入 email_content = str(email_content) # 调用检测函数 prediction = detect_phishing(email_content) # 注意:detect_phishing已返回单个预测值,无需再取索引[0] return jsonify({'prediction': prediction}) except Exception as e: app.logger.error(f'An error occurred: {str(e)}') app.logger.info(traceback.format_exc()) return jsonify({'error': f'An error occurred: {str(e)}'}), HTTPStatus.INTERNAL_SERVER_ERROR else: return jsonify({'error': 'Invalid method'}), HTTPStatus.METHOD_NOT_ALLOWED
2. 修复预测结果的索引错误
原代码中detect_phishing返回prediction[0](单个整数),但predict_phishing里又重复调用prediction[0],会导致TypeError: 'int' object is not subscriptable。上面的修改已去掉多余的索引调用。
3. 给检测函数增加输入防御
在detect_phishing里增加输入类型检查,确保传入的是字符串:
def detect_phishing(input_mail): # 确保输入为字符串类型 if not isinstance(input_mail, str): input_mail = str(input_mail) input_data_feature = feature_extraction.transform([input_mail]) prediction = model.predict(input_data_feature) return prediction[0]
4. 验证前端数据格式
确保前端发送的JSON符合预期,比如:
{ "email_content": "这是一封测试邮件的完整内容,用于判断是否为钓鱼邮件" }
如果前端需要拆分主题和正文发送,要和后端的拼接逻辑对应。
内容的提问来源于stack exchange,提问作者Kaze
相关产品推荐
相关产品推荐

