URL后缀随机变化时,如何实现法国股票列表的分页爬取?
解决动态随机分页URL的爬取问题
这种网站用随机后缀做分页参数,靠数字拼接URL肯定行不通,正确的做法是从当前页面的分页控件里提取下一页的真实链接,跟着页面提供的路径走。
修改后的实现思路
- 从第一页开始,每次请求完成后,解析页面里的「下一页」按钮链接
- 把这个链接作为下一次请求的URL,循环10次即可
- 额外加个User-Agent头,避免被网站直接拦截
调整后的代码
import requests from bs4 import BeautifulSoup links = [] current_url = "https://www.zonebourse.com/bourse/actions/Europe-3/France-51/" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} for _ in range(10): response = requests.get(current_url, headers=headers) if not response.ok: print(f"请求页面失败: {current_url}") break soup = BeautifulSoup(response.text, "lxml") # 提取股票链接(保留你原来的逻辑) tds = soup.find_all("td") for td in tds: a = td.find("a") if a and a.get("href", "").startswith("/cours/action/"): full_link = f"https://www.zonebourse.com{a['href']}" links.append(full_link) # 定位下一页链接(需匹配页面实际的分页按钮结构) next_page = soup.find("a", class_="pagination-next") if not next_page or not next_page.get("href"): print("找不到下一页链接,提前终止") break # 更新当前URL为下一页地址 current_url = f"https://www.zonebourse.com{next_page['href']}" print(f"共抓取到 {len(links)} 条股票链接")
注意事项
- 要手动确认页面里「下一页」按钮的HTML结构,比如class是否为
pagination-next,如果页面更新,需要同步调整find的参数 - 如果网站有反爬限制,可添加
time.sleep(1)这类延迟,或者使用代理避免IP被封 - 爬取前先打开目标页面,检查分页元素的具体属性,确保解析逻辑准确
内容的提问来源于stack exchange,提问作者baring
相关产品推荐
相关产品推荐

