You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BS4与Requests爬虫时所有页面返回相同数据的问题

问题描述

我用BS4和Requests写了个爬虫脚本,想抓取所有商品的URL,方便后续爬取商品信息。但遇到了问题:每个页面包含50个产品,运行脚本后列表里有300条数据(每个商品对应3个href),但集合里只有50个唯一商品,而且全是第一页的内容。不管用哪种分页方式,总是两次都拿到第一页的数据。

我试过用requests.Session()、修改请求头、添加请求间隔,都没解决问题。

我的代码

def all_URLs():
    list_urls = []
    url = ["https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11",
           "https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11#2"]

    with requests.Session() as s: # ADDED
        headers = {'User-Agent': "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/74.0.3729.169 YaBrowser/19.6.1.153 Yowser/2.5 Safari/537.36",
                    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
                    'Accept-Language': 'en-US,en;q=0.5',
                    'Accept-Encoding': 'gzip, deflate, br',
                    'Referer': 'https://www.gismeteo.ru/weather-orenburg-5159/now/',
                    'DNT': '1',
                    'Connection': 'false',
                    'Upgrade-Insecure-Requests': '1',
                    'Cache-Control': 'no-cache, max-age=0',
                    'TE': 'Trailers'} # ADDED
        for single_url in url:
            print(single_url)
            #try except to move on if any url is broken
            try:
                page = s.get(single_url,headers=headers)
                # Consider any status other than 2xx an error
                if not page.status_code // 100 == 2:
                    return "Error: Unexpected response {}".format(page)

            except requests.exceptions.RequestException as e:
                # A serious problem happened, like an SSLError or InvalidURL
                return "Error: {}".format(e)

            soup = BeautifulSoup(page.content, "html.parser")


            find_url = soup.find_all('a', class_='shelf-qd-v1-highlight', href=True)

            for url in find_url:
                #print(url)
                list_urls.append(url['href'])
            sleep(5)  # ADDED

    return list_urls

if __name__ == '__main__':

    all_urls = all_URLs()
    #get all  unique entries from the list
    set_of_urls = set(all_urls)
    print("X_X_X_X_X_X_X_")
    #print("list ",len(all_urls),"set ",len(set_of_urls))
    for a in set_of_urls:
        print(a)

运行输出

https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11
https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11#2
X_X_X_X_X_X_X_
list  300 set  50
Process finished with exit code 0

问题原因及解决方法

1. 分页参数错误

你用的第二个URL带了#2锚点,这是前端用来页面内定位的标记,服务器不会把它当作分页参数处理,所以请求这个URL时,服务器依然返回第一页内容。这类电商网站的分页通常用page、offset这类参数,比如你可以手动点击网页上的分页按钮,查看第二页的真实URL,把参数替换进去(比如可能是&page=2或者&offset=50)。

2. 变量名冲突

循环里的for url in find_url:会覆盖外层定义的url列表变量,虽然这次没直接导致分页问题,但会引发潜在bug,建议改成for item_url in find_url:。

3. 错误处理逻辑问题

当前错误处理用了return,一旦某个请求出错,整个函数会直接终止,后续URL不会继续处理。应该把return改成print加continue,这样遇到错误时跳过当前URL,继续处理下一个。

4. 请求头参数修正

  • Referer设置成了无关网站的链接,建议改成目标网站的域名(比如https://www.masrefacciones.mx/),避免被反爬策略拦截。
  • Connection的有效值是keep-alive或close,之前的false是错误值,可能影响请求稳定性。

修改后的代码示例

from bs4 import BeautifulSoup
import requests
from time import sleep

def all_URLs():
    list_urls = []
    # 修正分页参数,以实际网站分页按钮的URL为准
    url_list = ["https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11",
           "https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11&page=2"]

    with requests.Session() as s:
        headers = {'User-Agent': "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/74.0.3729.169 YaBrowser/19.6.1.153 Yowser/2.5 Safari/537.36",
                    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
                    'Accept-Language': 'en-US,en;q=0.5',
                    'Accept-Encoding': 'gzip, deflate, br',
                    'Referer': 'https://www.masrefacciones.mx/',
                    'DNT': '1',
                    'Connection': 'keep-alive',
                    'Upgrade-Insecure-Requests': '1',
                    'Cache-Control': 'no-cache, max-age=0',
                    'TE': 'Trailers'}
        for single_url in url_list:
            print(single_url)
            try:
                page = s.get(single_url, headers=headers)
                if not page.status_code // 100 == 2:
                    print(f"Error: Unexpected response {page.status_code} for {single_url}")
                    continue

            except requests.exceptions.RequestException as e:
                print(f"Error: {e} for {single_url}")
                continue

            soup = BeautifulSoup(page.content, "html.parser")
            find_url = soup.find_all('a', class_='shelf-qd-v1-highlight', href=True)

            for item_url in find_url:
                list_urls.append(item_url['href'])
            sleep(5)

    return list_urls

if __name__ == '__main__':
    all_urls = all_URLs()
    set_of_urls = set(all_urls)
    print("X_X_X_X_X_X_X_")
    print(f"list {len(all_urls)} set {len(set_of_urls)}")
    for a in set_of_urls:
        print(a)

内容的提问来源于stack exchange,提问作者BigBugNoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 11:58:10