使用BS4与Requests爬虫时所有页面返回相同数据的问题
问题描述
我用BS4和Requests写了个爬虫脚本,想抓取所有商品的URL,方便后续爬取商品信息。但遇到了问题:每个页面包含50个产品,运行脚本后列表里有300条数据(每个商品对应3个href),但集合里只有50个唯一商品,而且全是第一页的内容。不管用哪种分页方式,总是两次都拿到第一页的数据。
我试过用requests.Session()、修改请求头、添加请求间隔,都没解决问题。
我的代码
def all_URLs(): list_urls = [] url = ["https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11", "https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11#2"] with requests.Session() as s: # ADDED headers = {'User-Agent': "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/74.0.3729.169 YaBrowser/19.6.1.153 Yowser/2.5 Safari/537.36", 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://www.gismeteo.ru/weather-orenburg-5159/now/', 'DNT': '1', 'Connection': 'false', 'Upgrade-Insecure-Requests': '1', 'Cache-Control': 'no-cache, max-age=0', 'TE': 'Trailers'} # ADDED for single_url in url: print(single_url) #try except to move on if any url is broken try: page = s.get(single_url,headers=headers) # Consider any status other than 2xx an error if not page.status_code // 100 == 2: return "Error: Unexpected response {}".format(page) except requests.exceptions.RequestException as e: # A serious problem happened, like an SSLError or InvalidURL return "Error: {}".format(e) soup = BeautifulSoup(page.content, "html.parser") find_url = soup.find_all('a', class_='shelf-qd-v1-highlight', href=True) for url in find_url: #print(url) list_urls.append(url['href']) sleep(5) # ADDED return list_urls if __name__ == '__main__': all_urls = all_URLs() #get all unique entries from the list set_of_urls = set(all_urls) print("X_X_X_X_X_X_X_") #print("list ",len(all_urls),"set ",len(set_of_urls)) for a in set_of_urls: print(a)
运行输出
https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11 https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11#2 X_X_X_X_X_X_X_ list 300 set 50 Process finished with exit code 0
问题原因及解决方法
1. 分页参数错误
你用的第二个URL带了#2锚点,这是前端用来页面内定位的标记,服务器不会把它当作分页参数处理,所以请求这个URL时,服务器依然返回第一页内容。这类电商网站的分页通常用page、offset这类参数,比如你可以手动点击网页上的分页按钮,查看第二页的真实URL,把参数替换进去(比如可能是&page=2或者&offset=50)。
2. 变量名冲突
循环里的for url in find_url:会覆盖外层定义的url列表变量,虽然这次没直接导致分页问题,但会引发潜在bug,建议改成for item_url in find_url:。
3. 错误处理逻辑问题
当前错误处理用了return,一旦某个请求出错,整个函数会直接终止,后续URL不会继续处理。应该把return改成print加continue,这样遇到错误时跳过当前URL,继续处理下一个。
4. 请求头参数修正
Referer设置成了无关网站的链接,建议改成目标网站的域名(比如https://www.masrefacciones.mx/),避免被反爬策略拦截。Connection的有效值是keep-alive或close,之前的false是错误值,可能影响请求稳定性。
修改后的代码示例
from bs4 import BeautifulSoup import requests from time import sleep def all_URLs(): list_urls = [] # 修正分页参数,以实际网站分页按钮的URL为准 url_list = ["https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11", "https://www.masrefacciones.mx/tren-motriz/clutch-o-embrague/kit-de-clutch/valeo?PS=50&utmi_p=_tienda-oficial-valeo&utmi_pc=Banner%3acategoria11&utmi_cp=11&page=2"] with requests.Session() as s: headers = {'User-Agent': "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/74.0.3729.169 YaBrowser/19.6.1.153 Yowser/2.5 Safari/537.36", 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://www.masrefacciones.mx/', 'DNT': '1', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1', 'Cache-Control': 'no-cache, max-age=0', 'TE': 'Trailers'} for single_url in url_list: print(single_url) try: page = s.get(single_url, headers=headers) if not page.status_code // 100 == 2: print(f"Error: Unexpected response {page.status_code} for {single_url}") continue except requests.exceptions.RequestException as e: print(f"Error: {e} for {single_url}") continue soup = BeautifulSoup(page.content, "html.parser") find_url = soup.find_all('a', class_='shelf-qd-v1-highlight', href=True) for item_url in find_url: list_urls.append(item_url['href']) sleep(5) return list_urls if __name__ == '__main__': all_urls = all_URLs() set_of_urls = set(all_urls) print("X_X_X_X_X_X_X_") print(f"list {len(all_urls)} set {len(set_of_urls)}") for a in set_of_urls: print(a)
内容的提问来源于stack exchange,提问作者BigBugNoob
相关产品推荐
相关产品推荐

