使用Python爬取多页数据时后续页面返回首页结果的问题
问题原因及解决方法
你的问题出在URL参数格式错误,导致服务器无法正确识别page参数,始终返回第一页内容。
原URL里的sort=top_seller%3Fpage%3D5是错误的:%3F是问号?的URL编码,这会让page=5变成sort参数的一部分,而不是独立的分页参数。服务器只会识别到sort=top_seller?page=5,忽略后面重复的&page=5,所以始终返回第一页。
修正后的代码
from urllib.request import urlopen, Request from bs4 import BeautifulSoup # 正确的URL:用&分隔sort和page参数 url = 'https://tiki.vn/lam-sach-da-mat/c11232?sort=top_seller&page=5' # 添加请求头模拟浏览器,避免被反爬拦截 headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'} req = Request(url, headers=headers) html = urlopen(req) bs = BeautifulSoup(html, 'html.parser') # 找到所有product-item标签并提取product-id result = [tag.get('product-id') for tag in bs.find_all(lambda tag: tag.get('class') == ['product-item'])] print(result)
额外说明
- 直接用
urlopen发送请求可能被网站的反爬机制识别,返回默认页(第一页),所以添加User-Agent请求头模拟浏览器访问很有必要。 - 提取
product-id时,通过tag.get('product-id')可以直接获取标签的该属性值,得到你需要的结果。
内容的提问来源于stack exchange,提问作者Duchero Nguyễn
相关产品推荐
相关产品推荐

