Python提取含'attach'的网页链接时报错:NoneType不可迭代
提取含"attach"关键词的链接报错解决
问题情况
想用Python提取网页里包含“attach”的链接并存入列表,但运行代码时报错,错误指向第16行。
原代码
from bs4 import BeautifulSoup # pip install requests import requests def list_image_links(url): # Send a GET request to the URL response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") image_links = [] for image_links in soup.find_all('a'): href = image_links.get('href') #if attach is in href word = 'attach' if word in href: image_links.append(href) print(image_links) return image_links list_image_links('https://forum.ubuntu.ir/index.php?topic=211.0')
错误信息
Traceback (most recent call last): File "d:\...\main.py", line 22, in <module> list_image_links('https://forum.ubuntu.ir/index.php?topic=211.0') File "d:\...\main.py", line 16, in list_image_links if word in href: ^^^^^^^^^^^^ TypeError: argument of type 'NoneType' is not iterable
问题原因
- 变量名重复冲突:循环里用
image_links作为迭代变量,直接覆盖了之前定义的空列表,导致后续没法用append添加元素,逻辑彻底混乱。 - 未处理空值:页面里的
<a>标签不一定都有href属性,get('href')会返回None,直接判断字符串是否在None里,就触发了类型错误。
修复后的代码
from bs4 import BeautifulSoup import requests def list_image_links(url): response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") image_links = [] # 换个循环变量名,避免和结果列表重名 for link in soup.find_all('a'): href = link.get('href') # 先确认href不是空值,再检查关键词 if href and 'attach' in href: image_links.append(href) print(image_links) return image_links list_image_links('https://forum.ubuntu.ir/index.php?topic=211.0')
核心修复点
- 循环变量改用
link,避免和存储结果的image_links列表冲突。 - 增加
href非空判断,只有当href存在时,再检查是否包含“attach”关键词,解决NoneType报错问题。
内容的提问来源于stack exchange,提问作者Mr Hunter
相关产品推荐
相关产品推荐

