You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Drugbank.vn表格爬虫报错求助:AttributeError及多页爬取需求

修复爬虫代码并实现多页数据抓取

错误原因分析

报错AttributeError: 'NoneType' object has no attribute 'find_all'是因为soup.find('table')返回了None——服务器未返回包含目标表格的页面内容,大概率是请求缺少必要的请求头(比如User-Agent),被网站反爬机制拦截,返回了非目标页面。

修复后的完整代码

import pandas as pd
import requests
from bs4 import BeautifulSoup as bs

# 模拟浏览器的请求头,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

all_data = []
page_num = 1
page_size = 20

while True:
    # 构造当前分页的请求URL
    url = f"https://drugbank.vn/danh-sach/co-so-kinh-doanh?page={page_num}&size={page_size}&sort=id,desc"
    response = requests.get(url, headers=headers)
    
    # 检查请求是否成功
    if response.status_code != 200:
        print(f"请求第{page_num}页失败,状态码:{response.status_code}")
        break
    
    soup = bs(response.text, 'lxml')
    table = soup.find('table')
    
    # 找不到表格意味着没有更多数据,终止循环
    if not table:
        print(f"第{page_num}页未找到表格,停止抓取")
        break
    
    # 提取表格行数据,跳过表头行
    rows = table.find_all('tr')
    for row in rows[1:]:
        cols = row.find_all('td')
        cols = [ele.text.strip() for ele in cols]
        if cols:  # 跳过空行
            all_data.append(cols)
    
    print(f"已完成第{page_num}页数据抓取")
    page_num += 1

# 将抓取到的数据转为DataFrame并保存为CSV
df = pd.DataFrame(all_data)
# 可根据实际表格表头设置列名,示例:
# df.columns = ['ID', '经营主体名称', '地址', '联系电话', '状态']
print("全量数据抓取完成,预览前5行:")
print(df.head())
df.to_csv('drugstore_data.csv', index=False, encoding='utf-8-sig')

关键修复与优化说明

  • 添加User-Agent请求头:模拟真实浏览器访问,绕过基础反爬机制,确保能获取到包含表格的正常页面。
  • 循环分页抓取:通过递增page_num构造分页URL,直到找不到表格或请求失败,实现全量数据抓取。
  • 增加异常判断:检查请求状态码、表格是否存在,避免程序中途崩溃。
  • 过滤无效行:跳过表头行和空行,保证数据有效性。

内容的提问来源于stack exchange,提问作者Linh Tuấn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 09:02:31