多页网页爬取报错排查:Tableau客户页面爬虫代码修复
问题分析与解决
报错原因
报错是因为爬取到某一页时,info.find('img', {'class':'card__logo'})返回了None,直接访问['alt']触发了异常。主要有几个诱因:
- 多页请求被网站反爬,返回的页面内容异常,找不到目标元素
- 部分页面的元素结构和第一页不一致,没有对应的img标签
- 代码存在隐藏问题:
list_c在每一次页面循环里都重新初始化,最后只能保留最后一页的数据,前面爬取的内容全部丢失
修复后的代码
import requests from bs4 import BeautifulSoup import pandas as pd # 初始化存储所有数据的列表,放在页面循环外面 list_c = [] # 添加请求头,模拟浏览器访问,降低被反爬概率 headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } for n in range(0, 35): url = f'https://www.tableau.com/solutions/customers?region=All&industry=All&department=All&page={n}' # 带请求头发送请求 pages = requests.get(url, headers=headers) # 先检查请求是否成功 if pages.status_code != 200: print(f"第{n}页请求失败,状态码:{pages.status_code}") continue soup = BeautifulSoup(pages.text, 'html.parser') cus_info = soup.find_all('div', {'class': 'card__details'}) for info in cus_info: # 逐个元素判断是否存在,避免None时访问属性报错 logo_img = info.find('img', {'class':'card__logo'}) cust_name = logo_img['alt'] if logo_img else None cust_use_tag = info.find('h3', {'class':'card__title'}) cust_use = cust_use_tag.text.strip() if cust_use_tag else None cust_a_tag = info.find('a', {'class':''}) cust_url = 'https://www.tableau.com' + cust_a_tag['href'] if cust_a_tag else None cust_logo = 'https://www.tableau.com' + logo_img['src'] if logo_img else None customer_info = { 'cust_name': cust_name, 'cust_use': cust_use, 'cust_url': cust_url, 'cust_logo': cust_logo } list_c.append(customer_info) df_c = pd.DataFrame(list_c) print(df_c)
关键修改点
- 把
list_c移到页面循环外,确保所有页面的数据都能被存入 - 添加
headers模拟浏览器请求,降低被网站识别为爬虫的概率 - 每个元素查找后都做存在性判断,避免返回
None时访问属性报错 - 增加请求状态码检查,能快速定位请求失败的页面
- 拼接URL时补全
https://,避免生成无效链接
内容的提问来源于stack exchange,提问作者Luciol
相关产品推荐
相关产品推荐

