You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

URL后缀随机变化时,如何实现法国股票列表的分页爬取?

解决动态随机分页URL的爬取问题

这种网站用随机后缀做分页参数,靠数字拼接URL肯定行不通,正确的做法是从当前页面的分页控件里提取下一页的真实链接,跟着页面提供的路径走。

修改后的实现思路

  • 从第一页开始,每次请求完成后,解析页面里的「下一页」按钮链接
  • 把这个链接作为下一次请求的URL,循环10次即可
  • 额外加个User-Agent头,避免被网站直接拦截

调整后的代码

import requests
from bs4 import BeautifulSoup

links = []
current_url = "https://www.zonebourse.com/bourse/actions/Europe-3/France-51/"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}

for _ in range(10):
    response = requests.get(current_url, headers=headers)
    if not response.ok:
        print(f"请求页面失败: {current_url}")
        break
    
    soup = BeautifulSoup(response.text, "lxml")
    
    # 提取股票链接(保留你原来的逻辑)
    tds = soup.find_all("td")
    for td in tds:
        a = td.find("a")
        if a and a.get("href", "").startswith("/cours/action/"):
            full_link = f"https://www.zonebourse.com{a['href']}"
            links.append(full_link)
    
    # 定位下一页链接(需匹配页面实际的分页按钮结构)
    next_page = soup.find("a", class_="pagination-next")
    if not next_page or not next_page.get("href"):
        print("找不到下一页链接,提前终止")
        break
    
    # 更新当前URL为下一页地址
    current_url = f"https://www.zonebourse.com{next_page['href']}"

print(f"共抓取到 {len(links)} 条股票链接")

注意事项

  • 要手动确认页面里「下一页」按钮的HTML结构,比如class是否为pagination-next,如果页面更新,需要同步调整find的参数
  • 如果网站有反爬限制,可添加time.sleep(1)这类延迟,或者使用代理避免IP被封
  • 爬取前先打开目标页面,检查分页元素的具体属性,确保解析逻辑准确

内容的提问来源于stack exchange,提问作者baring

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 18:28:01