使用BeautifulSoup爬取Truecar多页数据时重复爬取第一页的解决方法
解决Truecar多页爬取重复第一页数据的问题
我尝试用Python的BeautifulSoup爬取Truecar网站的车辆信息,但当前代码在爬取下一页数据时,会重复爬取第一页的内容。请问如何实现无重复的多页数据爬取?以下是我的代码:
carname=input('Please enter the name of car: ') import requests from bs4 import BeautifulSoup import re import mysql.connector from selenium import webdriver from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager cnx = mysql.connector.connect(user='root',host='127.0.0.1',database='info',password='nooshika') cursor=cnx.cursor() listname=[] listprice=[] listtmile=[] listyearmodel=[] list2=[] final_list=[] #response=requests.get('https://www.truecar.com/used-cars-for-sale/listings/?page=1') #soup=BeautifulSoup(response.text,'html.parser') def scrape(response,soup): for link in soup.find_all('span',class_="truncate"): link=link.get_text() link=re.sub(r'Excellent Price',' ',str(link)) listname.append(link) for j in listname: if j==' ': listname.remove(j) for i in soup.find_all('div',class_="truncate text-xs",attrs={'data-test':"vehicleMileage"}): listtmile.append(i.get_text()) for k in soup.find_all('span',attrs={'data-test':"vehicleListingPriceAmount"}): k=k.get_text() listprice.append(k) for l in soup.find_all('span',class_="vehicle-card-year text-xs"): l=l.get_text() listyearmodel.append(l) for k,i2 in enumerate(listname): for k1,i3 in enumerate(listprice): for k2,i4 in enumerate(listtmile): for k3,i5 in enumerate(listyearmodel): if k==k1==k2==k3: list2.append((i5,i2,i3,i4)) pages = list(range(1,3)) for count in pages: response=requests.get('https://www.truecar.com/used-cars-for-sale/listings/?page={}'.format(count)) soup=BeautifulSoup(response.text,'html.parser') scrape(response,soup) def contains_carname(a): return isinstance(a, str) and carname in a matches = [x for x in list2 if contains_carname(x[1])] while matches==[]: matches = [x for x in list2 if contains_carname(x[1])] if matches!=[]: break print('No match.Please enter again') carname=input('Please enter the name of car: ') print(matches)
问题分析与修复方案
核心问题
- 全局列表累积重复数据:
listname、listprice等全局变量会持续追加所有页面的数据,再通过冗余的四重循环重新组合,导致list2出现大量重复项。 - 反爬拦截导致页面重复:Truecar会校验请求头,无UA的请求会被拦截,返回第一页内容,造成多页数据重复。
- 数据匹配逻辑低效且易出错:四重循环匹配索引完全冗余,车辆字段本身是一一对应的,无需多层遍历。
修复后的代码
carname = input('请输入车辆名称: ') import requests from bs4 import BeautifulSoup import re import mysql.connector # 数据库连接(保留原配置) cnx = mysql.connector.connect(user='root', host='127.0.0.1', database='info', password='nooshika') cursor = cnx.cursor() def scrape_page(page_num): # 模拟浏览器请求头,规避反爬 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } url = f'https://www.truecar.com/used-cars-for-sale/listings/?page={page_num}' response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 局部变量存储当前页数据,避免全局累积重复 current_page_data = [] # 先定位单个车辆卡片,确保字段一一对应 vehicle_cards = soup.find_all('div', class_='vehicle-card') for card in vehicle_cards: # 提取并清理车辆名称 name_elem = card.find('span', class_='truncate') name = name_elem.get_text(strip=True) if name_elem else '未知' name = re.sub(r'Excellent Price', '', name).strip() if not name: continue # 提取里程 mileage_elem = card.find('div', attrs={'data-test': 'vehicleMileage'}) mileage = mileage_elem.get_text(strip=True) if mileage_elem else '未知' # 提取价格 price_elem = card.find('span', attrs={'data-test': 'vehicleListingPriceAmount'}) price = price_elem.get_text(strip=True) if price_elem else '未知' # 提取年份 year_elem = card.find('span', class_='vehicle-card-year text-xs') year = year_elem.get_text(strip=True) if year_elem else '未知' current_page_data.append((year, name, price, mileage)) return current_page_data # 爬取第1-2页数据 final_list = [] for page in range(1, 3): page_data = scrape_page(page) final_list.extend(page_data) print(f"已爬取第{page}页,共{len(page_data)}条数据") # 筛选匹配车辆名称的结果(大小写不敏感) def contains_carname(item): return isinstance(item[1], str) and carname.lower() in item[1].lower() matches = [x for x in final_list if contains_carname(x)] while not matches: print('未找到匹配车辆,请重新输入') carname = input('请输入车辆名称: ') matches = [x for x in final_list if contains_carname(x)] print('匹配结果:') for item in matches: print(f"年份: {item[0]}, 名称: {item[1]}, 价格: {item[2]}, 里程: {item[3]}") # 可选:将数据存入数据库 # for item in matches: # cursor.execute("INSERT INTO cars (year, name, price, mileage) VALUES (%s, %s, %s, %s)", item) # cnx.commit() # 关闭数据库连接 cursor.close() cnx.close()
关键修改说明
- 局部化数据存储:将每页数据存储在函数局部变量中,处理完成后返回,避免全局列表累积重复数据。
- 添加请求头:模拟浏览器UA,解决反爬导致的页面重复问题。
- 优化数据提取逻辑:先定位单个车辆卡片,再从卡片内提取对应字段,确保数据一一对应,避免索引匹配错误。
- 简化筛选逻辑:改为大小写不敏感匹配,提升用户体验。
内容的提问来源于stack exchange,提问作者Ashley
相关产品推荐
相关产品推荐

