如何修复此场景下的TypeError: 'NoneType'对象不可迭代问题?
问题:遍历URL列表提取邮箱时触发NoneType错误
我编写脚本遍历URL列表查找邮箱地址,部分网站无结果返回,尝试了Stack Overflow上多种NoneType处理方案后仍报错。
导入的URL列表
url_list_updated = [ 'http://www.gfcadvice.com/', 'https://trillionfinancial.com.sg/about-us/', 'https://www.gen.com.sg/', 'https://www.aam-advisory.com/', 'https://www.proinvest.com.sg/', 'http://www.gilbertkoh.com/', 'https://dollarbureau.com/', 'http://www.greenfieldadvisory.com/', 'https://enpointefinancial.com/', 'https://www.ippfa.com/' ]
提取邮箱的原代码
for url in url_list_updated: response = requests.get(url) html_content = response.text soup = BeautifulSoup(html_content, 'html.parser') email_addresses = [] for link in soup.find_all('a'): # if 'mailto:' != None and 'mailto:' in link.get('href'): # if 'mailto:' != '' and 'mailto:' in link.get('href'): # if 'mailto:' in link.get('href') != None: if 'mailto:' in link.get('href') != '': email_addresses.append(link.get('href').replace('mailto:', '')) print(email_addresses) else: pass
报错信息
7 email_addresses = [] 8 for link in soup.find_all('a'): 9 # if 'mailto:' != None and 'mailto:' in link.get('href'): 10 # if 'mailto:' != '' and 'mailto:' in link.get('href'): 11 # if 'mailto:' in link.get('href') != None: ---> 12 if 'mailto:' in link.get('href') != '': 13 email_addresses.append(link.get('href').replace('mailto:', '')) 14 print(email_addresses) TypeError: argument of type 'NoneType' is not iterable
解决方案
核心问题是部分<a>标签没有href属性,调用link.get('href')会返回None,直接执行'mailto:' in None会触发TypeError。正确逻辑是先判断href是否存在且有效,再检查是否包含mailto:。
修改后的代码:
import requests from bs4 import BeautifulSoup url_list_updated = [ 'http://www.gfcadvice.com/', 'https://trillionfinancial.com.sg/about-us/', 'https://www.gen.com.sg/', 'https://www.aam-advisory.com/', 'https://www.proinvest.com.sg/', 'http://www.gilbertkoh.com/', 'https://dollarbureau.com/', 'http://www.greenfieldadvisory.com/', 'https://enpointefinancial.com/', 'https://www.ippfa.com/' ] for url in url_list_updated: try: response = requests.get(url, timeout=10) response.raise_for_status() # 捕获HTTP请求错误(如404、500) html_content = response.text soup = BeautifulSoup(html_content, 'html.parser') email_addresses = [] for link in soup.find_all('a'): href = link.get('href') # 先判断href不为None/空字符串,再检查是否包含mailto: if href and 'mailto:' in href: email = href.replace('mailto:', '') email_addresses.append(email) # 统一输出结果,避免重复打印 if email_addresses: print(f"从 {url} 提取到邮箱:{email_addresses}") else: print(f"未在 {url} 找到邮箱地址") except Exception as e: print(f"处理 {url} 时出错:{str(e)}")
关键修改点
- 先将
link.get('href')赋值给变量href,避免重复调用方法 - 用
if href同时处理None和空字符串的情况 - 增加
try-except块处理请求异常(如网站超时、无法访问等) - 调整打印逻辑,循环结束后统一输出结果,提升可读性
内容的提问来源于stack exchange,提问作者Haytorade
相关产品推荐
相关产品推荐

