You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Truecar多页数据时重复爬取第一页的解决方法

解决Truecar多页爬取重复第一页数据的问题

我尝试用Python的BeautifulSoup爬取Truecar网站的车辆信息,但当前代码在爬取下一页数据时,会重复爬取第一页的内容。请问如何实现无重复的多页数据爬取?以下是我的代码:

carname=input('Please enter the name of car: ')
import requests
from bs4 import BeautifulSoup
import re
import mysql.connector
from selenium import webdriver
from selenium.webdriver.common.by import By
from webdriver_manager.chrome import ChromeDriverManager 
cnx = mysql.connector.connect(user='root',host='127.0.0.1',database='info',password='nooshika')
cursor=cnx.cursor()
listname=[]
listprice=[]
listtmile=[]
listyearmodel=[]
list2=[]
final_list=[]
#response=requests.get('https://www.truecar.com/used-cars-for-sale/listings/?page=1')
#soup=BeautifulSoup(response.text,'html.parser')
def scrape(response,soup):
    for link in soup.find_all('span',class_="truncate"):
        link=link.get_text()
        link=re.sub(r'Excellent Price',' ',str(link))
        listname.append(link)
    for j in listname:
        if j==' ':
            listname.remove(j)
    for i in soup.find_all('div',class_="truncate text-xs",attrs={'data-test':"vehicleMileage"}):
        listtmile.append(i.get_text())
    for k in soup.find_all('span',attrs={'data-test':"vehicleListingPriceAmount"}):
        k=k.get_text()
        listprice.append(k)
    for l in soup.find_all('span',class_="vehicle-card-year text-xs"):
        l=l.get_text()
        listyearmodel.append(l)
    for k,i2 in enumerate(listname):
        for k1,i3 in enumerate(listprice):
            for k2,i4 in enumerate(listtmile):
                for k3,i5 in enumerate(listyearmodel):
                    if k==k1==k2==k3:
                        list2.append((i5,i2,i3,i4))

pages = list(range(1,3))
for count in pages:
    response=requests.get('https://www.truecar.com/used-cars-for-sale/listings/?page={}'.format(count))
    soup=BeautifulSoup(response.text,'html.parser')  
    scrape(response,soup)




def contains_carname(a):
    return isinstance(a, str) and carname in a
matches = [x for x in list2 if contains_carname(x[1])]
while matches==[]:
    matches = [x for x in list2 if contains_carname(x[1])]
    if matches!=[]:
        break
    print('No match.Please enter again')
    carname=input('Please enter the name of car: ')    
print(matches)

问题分析与修复方案

核心问题

  1. 全局列表累积重复数据:listname、listprice等全局变量会持续追加所有页面的数据,再通过冗余的四重循环重新组合,导致list2出现大量重复项。
  2. 反爬拦截导致页面重复:Truecar会校验请求头,无UA的请求会被拦截,返回第一页内容,造成多页数据重复。
  3. 数据匹配逻辑低效且易出错:四重循环匹配索引完全冗余,车辆字段本身是一一对应的,无需多层遍历。

修复后的代码

carname = input('请输入车辆名称: ')
import requests
from bs4 import BeautifulSoup
import re
import mysql.connector

# 数据库连接(保留原配置)
cnx = mysql.connector.connect(user='root', host='127.0.0.1', database='info', password='nooshika')
cursor = cnx.cursor()

def scrape_page(page_num):
    # 模拟浏览器请求头,规避反爬
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    url = f'https://www.truecar.com/used-cars-for-sale/listings/?page={page_num}'
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 局部变量存储当前页数据,避免全局累积重复
    current_page_data = []
    
    # 先定位单个车辆卡片,确保字段一一对应
    vehicle_cards = soup.find_all('div', class_='vehicle-card')
    for card in vehicle_cards:
        # 提取并清理车辆名称
        name_elem = card.find('span', class_='truncate')
        name = name_elem.get_text(strip=True) if name_elem else '未知'
        name = re.sub(r'Excellent Price', '', name).strip()
        if not name:
            continue
        
        # 提取里程
        mileage_elem = card.find('div', attrs={'data-test': 'vehicleMileage'})
        mileage = mileage_elem.get_text(strip=True) if mileage_elem else '未知'
        
        # 提取价格
        price_elem = card.find('span', attrs={'data-test': 'vehicleListingPriceAmount'})
        price = price_elem.get_text(strip=True) if price_elem else '未知'
        
        # 提取年份
        year_elem = card.find('span', class_='vehicle-card-year text-xs')
        year = year_elem.get_text(strip=True) if year_elem else '未知'
        
        current_page_data.append((year, name, price, mileage))
    
    return current_page_data

# 爬取第1-2页数据
final_list = []
for page in range(1, 3):
    page_data = scrape_page(page)
    final_list.extend(page_data)
    print(f"已爬取第{page}页,共{len(page_data)}条数据")

# 筛选匹配车辆名称的结果(大小写不敏感)
def contains_carname(item):
    return isinstance(item[1], str) and carname.lower() in item[1].lower()

matches = [x for x in final_list if contains_carname(x)]
while not matches:
    print('未找到匹配车辆,请重新输入')
    carname = input('请输入车辆名称: ')
    matches = [x for x in final_list if contains_carname(x)]

print('匹配结果:')
for item in matches:
    print(f"年份: {item[0]}, 名称: {item[1]}, 价格: {item[2]}, 里程: {item[3]}")

# 可选:将数据存入数据库
# for item in matches:
#     cursor.execute("INSERT INTO cars (year, name, price, mileage) VALUES (%s, %s, %s, %s)", item)
# cnx.commit()

# 关闭数据库连接
cursor.close()
cnx.close()

关键修改说明

  1. 局部化数据存储:将每页数据存储在函数局部变量中,处理完成后返回,避免全局列表累积重复数据。
  2. 添加请求头:模拟浏览器UA,解决反爬导致的页面重复问题。
  3. 优化数据提取逻辑:先定位单个车辆卡片,再从卡片内提取对应字段,确保数据一一对应,避免索引匹配错误。
  4. 简化筛选逻辑:改为大小写不敏感匹配,提升用户体验。

内容的提问来源于stack exchange,提问作者Ashley

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 21:14:59