You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python邮件爬虫报错:'builtin_function_or_method'对象无decode属性

错误原因分析
  • 核心触发点:代码中url = urls.popleft未添加括号,导致url变量存储的是popleft方法对象而非实际URL字符串。当urllib.parse.urlsplit(url)处理该方法对象时,内部逻辑尝试调用decode()方法,但方法对象无此属性,直接抛出AttributeError。
解决方案及完整修正代码

以下是修复所有问题后的代码,同时补充了必要的初始化和逻辑完善:

from collections import deque
import urllib.parse
import requests
import re
from bs4 import BeautifulSoup

scraped_url = set()
emails = set() 
# 初始化URL队列,替换为你要爬取的初始目标URL
urls = deque(['https://example.com'])

count = 0
try:
    while len(urls):
        count += 1
        if count == 20:
            break
        # 修正:调用popleft()方法获取实际URL
        url = urls.popleft()
        scraped_url.add(url)

        parts = urllib.parse.urlsplit(url)
        base_url = '{0.scheme}://{0.netloc}'.format(parts)
        
        path = url[:url.rfind('/')+1] if '/' in parts.path else url

        # 修正:格式化字符串的语法错误
        print('[%d] Processing %s' % (count, url))
        try:
           # 修正:传入url参数给requests.get()
           response = requests.get(url)
           # 新增:捕获HTTP状态码错误
           response.raise_for_status()
        except(requests.exceptions.MissingSchema, requests.exceptions.ConnectionError, requests.exceptions.HTTPError):
            continue

        new_emails = set(re.findall(r'[a-z0-9\.\-+_]+@[a-z0-9\.\-+_]+\.[a-z]+', response.text, re.I))
        # 修正:将爬取到的邮箱添加到全局集合
        emails.update(new_emails)
        soup = BeautifulSoup(response.text, features="lxml")

        for anchor in soup.find_all("a"):
            link = anchor.attrs['href'] if 'href' in anchor.attrs else ''
            if link.startswith('/'):
                link = base_url + link
            elif not link.startswith('http'):
                link = path + link 
            if link not in urls and link not in scraped_url:
                urls.append(link)
except KeyboardInterrupt:
    print('[-]closing')

for mail in emails:
    print(mail)
其他关键修正说明
  • 初始化urls为deque对象并传入初始URL,原代码因urls未定义会直接报错。
  • 修复print语句的格式化语法,将错误的count(count, url)改为正确的(count, url)。
  • 给requests.get()补充url参数,否则无法发起HTTP请求。
  • 添加response.raise_for_status()捕获404、500等HTTP错误,避免无效响应继续处理。
  • 将解析到的new_emails更新到全局emails集合,否则最终无法输出爬取到的邮箱。

内容的提问来源于stack exchange,提问作者sherlock

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 13:40:54