爬取free-proxy.cz时遇TypeError: 'NoneType' object is not iterable问题求助
问题:free-proxy.cz爬虫端口抓取报错TypeError
我正在编写free-proxy.cz的Socks5代理爬虫,端口抓取部分无法正常工作,运行代码后触发报错。以下是我的代码:
import requests from bs4 import BeautifulSoup import base64 urls = ['http://free-proxy.cz/en/proxylist/country/all/socks5/date/all', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/2', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/3', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/4', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/5', ] for url in urls: r = requests.get(url) soup = BeautifulSoup(r.text, 'html.parser') table = soup.find('table', {'id': 'proxy_list'}) for row in table.find('tbody').find_all('tr'): for ip in row.find('script'): text=base64.b64decode(ip[29:-2:]) for port in row.find('span', attrs='fport'): print(port.get_text()) #ipadd=print(prt.decode('utf-8')+':'+ports)
运行后报错信息:
Traceback (most recent call last): File "LOCATION\main.py", line 22, in <module> for port in row.find('span', attrs='fport'): TypeError: 'NoneType' object is not iterable 80 45554 1080 1080
问题原因
- None对象遍历错误:
row.find('span', attrs='fport')返回了None,说明当前遍历的表格行中没有class为fport的span元素,直接用for循环遍历None就触发了TypeError。 - 无效行干扰:表格里存在非代理数据的行(比如表头分隔行、空行),这些行没有端口对应的
fport标签,导致find方法返回None。 - IP解析逻辑冗余:
row.find('script')返回单个script元素,用for循环遍历它属于多余操作,直接提取script的文本内容即可处理。
修复方案
先判断端口元素是否存在再处理,同时优化IP解析逻辑:
import requests from bs4 import BeautifulSoup import base64 urls = ['http://free-proxy.cz/en/proxylist/country/all/socks5/date/all', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/2', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/3', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/4', 'http://free-proxy.cz/en/proxylist/country/all/socks5/date/all/5', ] for url in urls: r = requests.get(url) soup = BeautifulSoup(r.text, 'html.parser') table = soup.find('table', {'id': 'proxy_list'}) if not table: continue tbody = table.find('tbody') if not tbody: continue for row in tbody.find_all('tr'): # 处理IP script_tag = row.find('script') if script_tag: script_content = script_tag.string if script_content and len(script_content) > 31: ip_bytes = base64.b64decode(script_content[29:-2]) ip = ip_bytes.decode('utf-8') else: ip = None else: ip = None # 处理端口 port_span = row.find('span', class_='fport') if port_span: port = port_span.get_text(strip=True) if ip: print(f"{ip}:{port}")
内容的提问来源于stack exchange,提问作者Kouros
相关产品推荐
相关产品推荐

