Python去除价格字符串中的€并转为int以计算年均车价
问题与解决方案
问题
- 爬取二手车网站时,获取的价格数据带有€符号,需要转为int类型计算年均车价,但触发错误:
ValueError: invalid literal for int() with base 10: 'price',尝试方案未解决。 - 年份为字符串类型,是否需要转为int以便运算?
爬虫代码
import requests import pandas as pd from bs4 import BeautifulSoup url = "https://jammer.ie/used-cars?page={}&per-page=12" all_data = [] for page in range(1, 4): # <-- increase number of pages here soup = BeautifulSoup(requests.get(url.format(page)).text, "html.parser") for car in soup.select(".car"): info = car.select_one(".top-info").get_text(strip=True, separator="|") info = info.split("|") if len(info) == 4: make, model, year, price = info else: make, year, price = info model = "N/A" dealer_name = car.select_one(".dealer-name h6").get_text( strip=True, separator=" " ) address = car.select_one(".address").get_text(strip=True) features = {} for feature in car.select(".car--features li"): k = feature.img["src"].split("/")[-1].split(".")[0] v = feature.span.text features[f"feature_{k}"] = v all_data.append( { "make": make, "model": model, "year": year, "price": price, "dealer_name": dealer_name, "address": address, "url": "https://jammer.ie" + car.select_one("a[href*=vehicle]")["href"], **features, } ) df = pd.DataFrame(all_data) # prints sample data to screen: print(df.tail().to_markdown(index=False)) # saves all data to CSV df.to_csv("data.csv", index=False)
尝试的错误代码
df = pd.read_csv('data.csv', usecols= ['price','year']) print(type("price")) print(int("price"))
错误原因
你尝试的代码里int("price")是直接把字符串字面量"price"转为整数,这完全错误——应该处理DataFrame中price列的具体数值,而不是列名本身。另外爬取的价格带€符号,必须先清理符号才能转数值类型。
解决方案
1. 价格数据处理
方式一:爬取时直接处理
在爬虫代码中获取price后,立即清理符号并转int:
# 清理价格中的€和千分位逗号,转为int price_clean = price.replace('€', '').replace(',', '').strip() price_int = int(price_clean) if price_clean.isdigit() else None # 处理可能的无效值
将原代码中的"price": price替换为"price": price_int即可。
方式二:读取CSV后处理
如果已经保存了CSV,读取后批量处理:
import pandas as pd df = pd.read_csv('data.csv') # 清理price列的符号 df['price'] = df['price'].str.replace('€', '').str.replace(',', '').str.strip() # 转为数值类型,无效值转为NaN df['price'] = pd.to_numeric(df['price'], errors='coerce') # 如需int类型,填充NaN后转换 df['price'] = df['price'].fillna(0).astype(int)
2. 年份类型处理
如果需要用年份做数值运算(比如计算车龄:当前年份 - 车辆出厂年份),建议转为int类型:
爬取时处理
year_int = int(year) if year.isdigit() else None
替换原代码中的"year": year为"year": year_int。
读取CSV后处理
df['year'] = pd.to_numeric(df['year'], errors='coerce').fillna(0).astype(int)
修正后的完整爬虫代码
import requests import pandas as pd from bs4 import BeautifulSoup url = "https://jammer.ie/used-cars?page={}&per-page=12" all_data = [] for page in range(1, 4): # <-- increase number of pages here soup = BeautifulSoup(requests.get(url.format(page)).text, "html.parser") for car in soup.select(".car"): info = car.select_one(".top-info").get_text(strip=True, separator="|") info = info.split("|") if len(info) == 4: make, model, year, price = info else: make, year, price = info model = "N/A" # 处理价格 price_clean = price.replace('€', '').replace(',', '').strip() price_int = int(price_clean) if price_clean.isdigit() else None # 处理年份 year_int = int(year) if year.isdigit() else None dealer_name = car.select_one(".dealer-name h6").get_text( strip=True, separator=" " ) address = car.select_one(".address").get_text(strip=True) features = {} for feature in car.select(".car--features li"): k = feature.img["src"].split("/")[-1].split(".")[0] v = feature.span.text features[f"feature_{k}"] = v all_data.append( { "make": make, "model": model, "year": year_int, "price": price_int, "dealer_name": dealer_name, "address": address, "url": "https://jammer.ie" + car.select_one("a[href*=vehicle]")["href"], **features, } ) df = pd.DataFrame(all_data) # prints sample data to screen: print(df.tail().to_markdown(index=False)) # saves all data to CSV df.to_csv("data.csv", index=False)
内容的提问来源于stack exchange,提问作者user20637309
相关产品推荐
相关产品推荐

