You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+BeautifulSoup爬取数据时去除字符串多余字符问题

Fixing Your Scraper's Specs Formatting Issue

Hey Ben, nice start on your Python scraper—sounds like you're almost there! The issue you're seeing with [] and '' in your output is because you're slicing the specs list instead of grabbing individual elements directly. Let's break this down and fix it.

What's Going Wrong?

When you use slicing like specs[1:2], you're getting a sub-list containing one element, not the element itself. That's why your output looks like ['Petrol'] instead of just Petrol. The 'list' object has no attribute 'text' error happens because you tried using .text on a list (which doesn't have that method)—you only need .text for BeautifulSoup elements, not for strings already extracted into a list.

The Fix

Instead of slicing, access the specific index of the element you want from the specs list. Based on your current output, here's how to adjust those lines:

Replace this:

gearbox = specs[3:]
fuel = specs[1:2]
mileage = specs[0:1]

With this:

# Grab individual elements from the list instead of slices
gearbox = specs[3]
fuel = specs[1]
mileage = specs[0]

This will pull the raw string values directly, so your output will be Manual, Petrol, 86,863 miles as expected—perfect for writing clean CSV columns.

Modified Full Code

Here's your complete code with the fix applied:

import csv
import requests
from bs4 import BeautifulSoup

outfile = open('pistonheads.csv','w', newline='')
writer = csv.writer(outfile)
writer.writerow(["Link", "Make", "Model", "Price", "Image Link", "Gearbox", "Fuel", "Mileage"])

url = 'https://www.pistonheads.com/classifieds?Category=used-cars&Page=1&ResultsPerPage=100'
get_url = requests.get(url)
get_text = get_url.text
soup = BeautifulSoup(get_text, 'html.parser')

car_link = soup.find_all('div', 'listing-headline', 'price')
for div in car_link:
    links = div.findAll('a')
    for a in links:
        link = "https://www.pistonheads.com" + a['href']
        make = a['href'].split('/')[-4]
        model = a['href'].split('/')[-3]
        price = a.find('span').text.rstrip()
        
        image_link = a.parent.parent.find('img')['src']
        image = "https:" + image_link
        
        vehicle_details = a.parent.parent.find('ul', class_='specs')
        specs = list(vehicle_details.stripped_strings)
        
        # Fixed: Grab individual elements instead of slices
        gearbox = specs[3]
        fuel = specs[1]
        mileage = specs[0]
        
        writer.writerow([link, make, model, price, image, gearbox, fuel, mileage])
        print(link, make, model, price, image, gearbox, fuel, mileage)

outfile.close()

Quick Note

Double-check the indices in specs to make sure they match the actual order of data on the page. If the specs ever change order, you might need to adjust these numbers—but based on your current output, this should work perfectly.

内容的提问来源于stack exchange,提问作者Ben Willis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:21:11