Python网页爬虫生成CSV格式化需求:去空行、加指定表头及清理内容
Hey there! Let's tweak your code to meet all three of your CSV requirements—no more blank rows, the right headers, and no brackets cluttering up the car names. Here's how we'll do it step by step:
1. Remove Blank Rows
Extra blank lines often pop up when using the csv module without specifying newline='' when opening files. We'll fix this, plus use a with statement to handle file operations safely (it auto-closes the file for you, so no more forgetting close()!).
2. Add the Correct Headers
We'll replace the old ['brand_Names', 'Prices'] header with your requested columns: "Car make", "Car Model", "Price".
3. Get Rid of [] in Car Names
The brackets are showing up because you were writing a list (full_Names) directly to the CSV. Instead, we'll extract make and model as separate, clean strings, then write them as individual columns.
Updated Code
import requests from bs4 import BeautifulSoup as bs import csv import re source = requests.get('website link goes here...').text soup = bs(source, 'html.parser') # Separate lists for make, model, and price to keep data organized car_makes = [] car_models = [] car_prices = [] # Scrape car makes and models for var1 in soup.find_all('h3', class_='brandModelTitle'): car_names = var1.text test_split = car_names.split("\n") # Extract and clean make/model (strip extra whitespace) make = test_split[1].strip() model = test_split[2].strip() car_makes.append(make) car_models.append(model) # Scrape and clean prices for price_tag in soup.find_all('span', class_='f20 bold fieldPrice'): price = price_tag.span.text # Remove leading/trailing whitespace from prices price = re.sub("^\s+|\s+$", "", price, flags=re.UNICODE) car_prices.append(price) # Write final CSV with all requirements with open('pyScraper_1.3_Export.csv', 'w', newline='', encoding='utf-8') as csv_file: csv_writer = csv.writer(csv_file) # Write the requested headers csv_writer.writerow(['Car make', 'Car Model', 'Price']) # Loop through matching data and write each row for make, model, price in zip(car_makes, car_models, car_prices): csv_writer.writerow([make, model, price])
Quick Breakdown of Key Changes
- No more brackets: By pulling make and model as individual stripped strings instead of a list, they write to the CSV without
[]. - No blank rows: The
newline=''parameter inopen()stops extra empty lines from appearing between rows. - Clean headers: The first row now uses exactly the column names you wanted.
- Safer file handling: The
withstatement takes care of closing the file automatically, so you don't have to worry about missedclose()calls.
If you ever run into a case where the number of makes/models doesn't match the number of prices, you can add a quick check to ensure all lists are the same length before writing—but assuming your scrape collects one price per car, this should work perfectly.
内容的提问来源于stack exchange,提问作者Dython

