如何用Python提取ul下指定li标签并导出至Excel多列
解决方案
1. 提取ul标签下的特定li元素内容
Autotrader产品卡片的关键信息(年份、里程、车身类型等)都在class="product-card-key-facts__list"的ul标签下,我们可以通过遍历li元素、匹配文本特征来提取对应字段:
修改后的完整代码
from requests_html import HTMLSession from bs4 import BeautifulSoup import csv # 如需直接导出Excel,先安装pandas后取消以下注释 # import pandas as pd s = HTMLSession() url = 'https://www.autotrader.co.uk/car-search?postcode=CV23%208AJ&include-delivery-option=on&advertising-location=at_cars&page=1' page = s.get(url) soup = BeautifulSoup(page.content, 'html.parser') product_cards = soup.find_all('div', class_="product-card__inner") # 定义表头与数据存储列表 header = ['Price', 'Title', 'Subtitle', 'Year', 'Body', 'Miles', 'Engine', 'Power', 'Gearbox', 'Fuel', 'Owners', 'Ulez', 'Other1', 'Other2', 'Other3'] all_data = [] for card in product_cards: # 提取基础信息 price = card.find('div', class_="product-card-pricing__price").get_text(strip=True) title = card.find('h3', class_="product-card-details__title").get_text(strip=True) subtitle = card.find('p', class_="product-card-details__subtitle").get_text(strip=True) # 初始化关键信息字段 year = body = miles = engine = power = gearbox = fuel = owners = ulez = other1 = other2 = other3 = "" # 定位关键信息的ul列表 key_facts_list = card.find('ul', class_='product-card-key-facts__list') if key_facts_list: key_facts = key_facts_list.find_all('li', class_='product-card-key-facts__item') # 遍历li元素,根据文本特征匹配对应字段 for idx, li in enumerate(key_facts): text = li.get_text(strip=True) if 'reg' in text: year = text elif any(kw in text for kw in ['dr', 'Hatchback', 'Saloon', 'Estate', 'SUV', 'Coupe']): body = text elif 'miles' in text: miles = text elif 'cc' in text: engine = text elif 'bhp' in text: power = text elif any(kw in text for kw in ['Manual', 'Automatic', 'Semi-Auto']): gearbox = text elif any(kw in text for kw in ['Petrol', 'Diesel', 'Electric', 'Hybrid']): fuel = text elif 'Owner' in text: owners = text elif 'ULEZ' in text: ulez = text # 处理未匹配的li元素,按顺序存入备用字段 else: if idx == 0: other1 = text elif idx == 1: other2 = text elif idx == 2: other3 = text # 组装当前卡片的完整数据 row_data = [price, title, subtitle, year, body, miles, engine, power, gearbox, fuel, owners, ulez, other1, other2, other3] all_data.append(row_data) # 导出为CSV文件(可直接用Excel打开,字段自动对应列) with open('autotraderdata.csv', 'w', encoding='utf-8', newline='') as f: writer = csv.writer(f) writer.writerow(header) writer.writerows(all_data) # 如需直接导出Excel格式,取消以下注释(需先执行:pip install pandas openpyxl) # df = pd.DataFrame(all_data, columns=header) # df.to_excel('autotraderdata.xlsx', index=False, engine='openpyxl') print("数据导出完成")
2. 导出至Excel不同列
- 上述代码生成的CSV文件可以直接用Microsoft Excel打开,每个字段会自动对应到独立列;
- 若需要直接生成
.xlsx格式文件,安装pandas和openpyxl库后,取消代码中对应注释即可,导出的Excel文件会严格按照表头将数据分配到不同列。
注意事项
- 字段匹配逻辑基于当前Autotrader页面的文本特征编写,若网站结构或文本表述变更,需对应调整判断条件;
- 未匹配到预设字段的li元素会存入
Other1/Other2/Other3列,可根据实际需求修改这部分逻辑。
内容的提问来源于stack exchange,提问作者Tarek
相关产品推荐
相关产品推荐

