Python+BeautifulSoup爬取数据时去除字符串多余字符问题
Hey Ben, nice start on your Python scraper—sounds like you're almost there! The issue you're seeing with [] and '' in your output is because you're slicing the specs list instead of grabbing individual elements directly. Let's break this down and fix it.
What's Going Wrong?
When you use slicing like specs[1:2], you're getting a sub-list containing one element, not the element itself. That's why your output looks like ['Petrol'] instead of just Petrol. The 'list' object has no attribute 'text' error happens because you tried using .text on a list (which doesn't have that method)—you only need .text for BeautifulSoup elements, not for strings already extracted into a list.
The Fix
Instead of slicing, access the specific index of the element you want from the specs list. Based on your current output, here's how to adjust those lines:
Replace this:
gearbox = specs[3:] fuel = specs[1:2] mileage = specs[0:1]
With this:
# Grab individual elements from the list instead of slices gearbox = specs[3] fuel = specs[1] mileage = specs[0]
This will pull the raw string values directly, so your output will be Manual, Petrol, 86,863 miles as expected—perfect for writing clean CSV columns.
Modified Full Code
Here's your complete code with the fix applied:
import csv import requests from bs4 import BeautifulSoup outfile = open('pistonheads.csv','w', newline='') writer = csv.writer(outfile) writer.writerow(["Link", "Make", "Model", "Price", "Image Link", "Gearbox", "Fuel", "Mileage"]) url = 'https://www.pistonheads.com/classifieds?Category=used-cars&Page=1&ResultsPerPage=100' get_url = requests.get(url) get_text = get_url.text soup = BeautifulSoup(get_text, 'html.parser') car_link = soup.find_all('div', 'listing-headline', 'price') for div in car_link: links = div.findAll('a') for a in links: link = "https://www.pistonheads.com" + a['href'] make = a['href'].split('/')[-4] model = a['href'].split('/')[-3] price = a.find('span').text.rstrip() image_link = a.parent.parent.find('img')['src'] image = "https:" + image_link vehicle_details = a.parent.parent.find('ul', class_='specs') specs = list(vehicle_details.stripped_strings) # Fixed: Grab individual elements instead of slices gearbox = specs[3] fuel = specs[1] mileage = specs[0] writer.writerow([link, make, model, price, image, gearbox, fuel, mileage]) print(link, make, model, price, image, gearbox, fuel, mileage) outfile.close()
Quick Note
Double-check the indices in specs to make sure they match the actual order of data on the page. If the specs ever change order, you might need to adjust these numbers—but based on your current output, this should work perfectly.
内容的提问来源于stack exchange,提问作者Ben Willis

