求助:用Python+BeautifulSoup爬取商品价格和日期的代码无法执行
亚马逊商品爬取代码排查与修复
问题描述
作为Python初学者,尝试用Python+BeautifulSoup爬取亚马逊商品的标题、价格并记录日期,代码无报错但无法执行,原代码如下:
from bs4 import BeautifulSoup import requests import time import datetime import pandas as pd import csv URL = 'https://www.amazon.com/Atomic-Habits-James-Clear-audiobook/dp/B07RFSSYBH/ref=sr_1_1?keywords=atomic+habits&qid=1685192621&s=books&sprefix=atomi%2Cstripbooks-intl-ship%2C315&sr=1-1' HEADERS = ({'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36','Accept-Language': 'en-US, en:q =0.5' }) page = requests.get(URL, headers = HEADERS) soup1 = BeautifulSoup(page.content, "html.parser") soup2 = BeautifulSoup(soup1.prettify(),"html.parser") title = soup2.find(id = 'product Title').get_text() price = soup2.find("span", attrs ={"class": 'a-size-base a-color-secondary'}).text print(title) print(price) price = price.strip()[1:] title = title.strip() print(price) print(title) today = datetime.date.today() print(today) # inserting data into excel header = ['title', 'price', 'Date'] data = [title, price, today] with open('AmazonWebScraper.csv', 'w', newline='',encoding='UTF8') as f: writer = csv.writer(f) writer.writerow(header) writer.writerow(data) df = pd.read_csv(r'C:\Users\User\Desktop\AmazonWebScraper.csv') print(df)
问题排查
- 标题ID错误:原代码中
id='product Title'包含空格,亚马逊实际标题的ID是productTitle(无空格),导致无法定位元素。 - 价格选择器错误:选择的
a-size-base a-color-secondary是辅助文本类,并非商品价格元素,该有声书的价格实际位于#audible_product_price节点下。 - 冗余HTML解析:
soup2 = BeautifulSoup(soup1.prettify(),"html.parser")属于多余操作,会改变HTML结构,增加解析失败概率,直接使用soup1即可。 - 请求头格式错误:
Accept-Language字段中en:q =0.5语法有误,应改为en;q=0.5(分号分隔,无空格),否则可能被亚马逊反爬机制拦截。 - 文件路径不一致:写入CSV用相对路径,读取却用绝对路径,若脚本不在桌面会导致读取失败,建议统一路径。
修复后的代码
from bs4 import BeautifulSoup import requests import datetime import pandas as pd import csv URL = 'https://www.amazon.com/Atomic-Habits-James-Clear-audiobook/dp/B07RFSSYBH/ref=sr_1_1?keywords=atomic+habits&qid=1685192621&s=books&sprefix=atomi%2Cstripbooks-intl-ship%2C315&sr=1-1' # 修正请求头格式 HEADERS = { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/113.0.0.0 Safari/537.36', 'Accept-Language': 'en-US, en;q=0.5' } # 获取并解析页面内容 page = requests.get(URL, headers=HEADERS) soup = BeautifulSoup(page.content, "html.parser") # 提取标题(修正ID) title = soup.find(id='productTitle').get_text(strip=True) # 提取价格(修正选择器,增加备用方案) price_element = soup.find("span", id="audible_product_price") if price_element: price = price_element.get_text(strip=True).replace('$', '') else: price_element = soup.find("span", class_="a-price-whole") price = price_element.get_text(strip=True) if price_element else "价格未获取" print("标题:", title) print("价格:", price) today = datetime.date.today() print("日期:", today) # 统一路径操作CSV csv_path = 'AmazonWebScraper.csv' header = ['title', 'price', 'Date'] data = [title, price, today] with open(csv_path, 'w', newline='', encoding='UTF8') as f: writer = csv.writer(f) writer.writerow(header) writer.writerow(data) df = pd.read_csv(csv_path) print(df)
额外提示
- 亚马逊反爬机制严格,若仍无法获取内容,可尝试从浏览器复制
cookies添加到请求头,或增加请求间隔(比如time.sleep(2))。 - 页面结构可能随时间变化,若后续失效,需重新查看页面HTML源码调整选择器。
内容的提问来源于stack exchange,提问作者ilda cengu
相关产品推荐
相关产品推荐

