You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python抓取多页面网站及指定网页表格?

Python实现多页面抓取及指定页面表格提取

所需依赖

先装必要的库:

pip install requests beautifulsoup4 pandas

第一步:抓取指定页面的表格

目标链接的表格是标准HTML table结构,用pandas.read_html可以直接提取,比BeautifulSoup更高效:

import requests
import pandas as pd

# 请求头,模拟浏览器访问,避免被反爬
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

target_url = 'https://haexpeditions.com/advice/list-of-mount-everest-climbers/'

# 发送请求获取页面内容
response = requests.get(target_url, headers=headers)
response.raise_for_status()  # 检查请求是否成功

# 用pandas直接读取页面中的所有表格
tables = pd.read_html(response.text)
# 通常第一个表格就是目标(如果页面有多个表格,可通过索引或筛选列名定位)
everest_climbers_df = tables[0]

# 保存为CSV文件
everest_climbers_df.to_csv('everest_climbers.csv', index=False)
print("表格抓取完成,已保存为everest_climbers.csv")

如果想用BeautifulSoup手动解析(适合需要自定义处理的场景):

from bs4 import BeautifulSoup

soup = BeautifulSoup(response.text, 'html.parser')
table = soup.find('table')  # 找到页面中的第一个表格

# 提取表头
headers = [th.text.strip() for th in table.find('thead').find_all('th')]
# 提取表格行数据
rows = []
for tr in table.find('tbody').find_all('tr'):
    row_data = [td.text.strip() for td in tr.find_all('td')]
    rows.append(row_data)

# 转为DataFrame
df = pd.DataFrame(rows, columns=headers)
df.to_csv('everest_climbers_bs4.csv', index=False)

第二步:多页面抓取的通用实现

多页面抓取核心是找到分页规律,常见的分页类型有两种:

1. URL带分页参数(比如?page=1、?p=2)

假设目标网站分页URL格式为https://example.com/list?page=数字,代码示例:

import requests
import pandas as pd
import time

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

base_url = 'https://example.com/list?page={}'
max_pages = 5  # 假设要爬前5页
all_data = []

for page_num in range(1, max_pages + 1):
    url = base_url.format(page_num)
    try:
        response = requests.get(url, headers=headers)
        response.raise_for_status()
        
        # 提取当前页数据(这里以抓取表格为例,可根据需求替换成其他元素提取逻辑)
        tables = pd.read_html(response.text)
        page_df = tables[0]
        all_data.append(page_df)
        
        print(f"第{page_num}页抓取完成")
        time.sleep(1)  # 控制请求频率,避免被封
    except Exception as e:
        print(f"第{page_num}页抓取失败:{str(e)}")
        continue

# 合并所有页面数据
final_df = pd.concat(all_data, ignore_index=True)
final_df.to_csv('multi_page_data.csv', index=False)
print("多页面抓取完成,已保存为multi_page_data.csv")

2. 通过“下一页”链接跳转

如果分页没有明显的数字参数,需要从页面中提取“下一页”的链接:

from bs4 import BeautifulSoup
import time

current_url = 'https://example.com/list'
all_data = []

while current_url:
    try:
        response = requests.get(current_url, headers=headers)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 提取当前页数据
        table = soup.find('table')
        headers = [th.text.strip() for th in table.find('thead').find_all('th')]
        rows = [[td.text.strip() for td in tr.find_all('td')] for tr in table.find('tbody').find_all('tr')]
        page_df = pd.DataFrame(rows, columns=headers)
        all_data.append(page_df)
        
        # 查找下一页链接(根据实际页面的“下一页”文本或class调整)
        next_link = soup.find('a', {'class': 'next-page'}) or soup.find('a', text='Next')
        if next_link and 'href' in next_link.attrs:
            current_url = next_link['href']
            # 如果是相对链接,拼接成绝对URL
            if not current_url.startswith('http'):
                current_url = 'https://example.com' + current_url
        else:
            current_url = None  # 没有下一页,结束循环
            
        print(f"当前页抓取完成,下一页链接:{current_url}")
        time.sleep(1)
    except Exception as e:
        print(f"抓取失败:{str(e)}")
        break

final_df = pd.concat(all_data, ignore_index=True)
final_df.to_csv('multi_page_data.csv', index=False)

注意事项

  • 必须带User-Agent请求头,大部分网站会拦截无标识的请求
  • 加入time.sleep()控制请求频率,避免触发反爬机制
  • 异常处理要到位,捕获请求失败、元素找不到等情况,保证程序稳定

内容的提问来源于stack exchange,提问作者juststuck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:50:25