Python邮件爬虫报错:'builtin_function_or_method'对象无decode属性
错误原因分析
- 核心触发点:代码中
url = urls.popleft未添加括号,导致url变量存储的是popleft方法对象而非实际URL字符串。当urllib.parse.urlsplit(url)处理该方法对象时,内部逻辑尝试调用decode()方法,但方法对象无此属性,直接抛出AttributeError。
解决方案及完整修正代码
以下是修复所有问题后的代码,同时补充了必要的初始化和逻辑完善:
from collections import deque import urllib.parse import requests import re from bs4 import BeautifulSoup scraped_url = set() emails = set() # 初始化URL队列,替换为你要爬取的初始目标URL urls = deque(['https://example.com']) count = 0 try: while len(urls): count += 1 if count == 20: break # 修正:调用popleft()方法获取实际URL url = urls.popleft() scraped_url.add(url) parts = urllib.parse.urlsplit(url) base_url = '{0.scheme}://{0.netloc}'.format(parts) path = url[:url.rfind('/')+1] if '/' in parts.path else url # 修正:格式化字符串的语法错误 print('[%d] Processing %s' % (count, url)) try: # 修正:传入url参数给requests.get() response = requests.get(url) # 新增:捕获HTTP状态码错误 response.raise_for_status() except(requests.exceptions.MissingSchema, requests.exceptions.ConnectionError, requests.exceptions.HTTPError): continue new_emails = set(re.findall(r'[a-z0-9\.\-+_]+@[a-z0-9\.\-+_]+\.[a-z]+', response.text, re.I)) # 修正:将爬取到的邮箱添加到全局集合 emails.update(new_emails) soup = BeautifulSoup(response.text, features="lxml") for anchor in soup.find_all("a"): link = anchor.attrs['href'] if 'href' in anchor.attrs else '' if link.startswith('/'): link = base_url + link elif not link.startswith('http'): link = path + link if link not in urls and link not in scraped_url: urls.append(link) except KeyboardInterrupt: print('[-]closing') for mail in emails: print(mail)
其他关键修正说明
- 初始化
urls为deque对象并传入初始URL,原代码因urls未定义会直接报错。 - 修复
print语句的格式化语法,将错误的count(count, url)改为正确的(count, url)。 - 给
requests.get()补充url参数,否则无法发起HTTP请求。 - 添加
response.raise_for_status()捕获404、500等HTTP错误,避免无效响应继续处理。 - 将解析到的
new_emails更新到全局emails集合,否则最终无法输出爬取到的邮箱。
内容的提问来源于stack exchange,提问作者sherlock
相关产品推荐
相关产品推荐

