You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取ul下指定li标签并导出至Excel多列

解决方案

1. 提取ul标签下的特定li元素内容

Autotrader产品卡片的关键信息(年份、里程、车身类型等)都在class="product-card-key-facts__list"的ul标签下,我们可以通过遍历li元素、匹配文本特征来提取对应字段:

修改后的完整代码

from requests_html import HTMLSession
from bs4 import BeautifulSoup
import csv
# 如需直接导出Excel,先安装pandas后取消以下注释
# import pandas as pd

s = HTMLSession()
url = 'https://www.autotrader.co.uk/car-search?postcode=CV23%208AJ&include-delivery-option=on&advertising-location=at_cars&page=1'
page = s.get(url)

soup = BeautifulSoup(page.content, 'html.parser')
product_cards = soup.find_all('div', class_="product-card__inner")

# 定义表头与数据存储列表
header = ['Price', 'Title', 'Subtitle', 'Year', 'Body', 'Miles', 'Engine', 'Power', 'Gearbox', 'Fuel', 'Owners', 'Ulez', 'Other1', 'Other2', 'Other3']
all_data = []

for card in product_cards:
    # 提取基础信息
    price = card.find('div', class_="product-card-pricing__price").get_text(strip=True)
    title = card.find('h3', class_="product-card-details__title").get_text(strip=True)
    subtitle = card.find('p', class_="product-card-details__subtitle").get_text(strip=True)
    
    # 初始化关键信息字段
    year = body = miles = engine = power = gearbox = fuel = owners = ulez = other1 = other2 = other3 = ""
    
    # 定位关键信息的ul列表
    key_facts_list = card.find('ul', class_='product-card-key-facts__list')
    if key_facts_list:
        key_facts = key_facts_list.find_all('li', class_='product-card-key-facts__item')
        # 遍历li元素,根据文本特征匹配对应字段
        for idx, li in enumerate(key_facts):
            text = li.get_text(strip=True)
            if 'reg' in text:
                year = text
            elif any(kw in text for kw in ['dr', 'Hatchback', 'Saloon', 'Estate', 'SUV', 'Coupe']):
                body = text
            elif 'miles' in text:
                miles = text
            elif 'cc' in text:
                engine = text
            elif 'bhp' in text:
                power = text
            elif any(kw in text for kw in ['Manual', 'Automatic', 'Semi-Auto']):
                gearbox = text
            elif any(kw in text for kw in ['Petrol', 'Diesel', 'Electric', 'Hybrid']):
                fuel = text
            elif 'Owner' in text:
                owners = text
            elif 'ULEZ' in text:
                ulez = text
            # 处理未匹配的li元素,按顺序存入备用字段
            else:
                if idx == 0:
                    other1 = text
                elif idx == 1:
                    other2 = text
                elif idx == 2:
                    other3 = text
    
    # 组装当前卡片的完整数据
    row_data = [price, title, subtitle, year, body, miles, engine, power, gearbox, fuel, owners, ulez, other1, other2, other3]
    all_data.append(row_data)

# 导出为CSV文件(可直接用Excel打开,字段自动对应列)
with open('autotraderdata.csv', 'w', encoding='utf-8', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(header)
    writer.writerows(all_data)

# 如需直接导出Excel格式,取消以下注释(需先执行:pip install pandas openpyxl)
# df = pd.DataFrame(all_data, columns=header)
# df.to_excel('autotraderdata.xlsx', index=False, engine='openpyxl')

print("数据导出完成")

2. 导出至Excel不同列

  • 上述代码生成的CSV文件可以直接用Microsoft Excel打开,每个字段会自动对应到独立列;
  • 若需要直接生成.xlsx格式文件,安装pandas和openpyxl库后,取消代码中对应注释即可,导出的Excel文件会严格按照表头将数据分配到不同列。

注意事项

  • 字段匹配逻辑基于当前Autotrader页面的文本特征编写,若网站结构或文本表述变更,需对应调整判断条件;
  • 未匹配到预设字段的li元素会存入Other1/Other2/Other3列,可根据实际需求修改这部分逻辑。

内容的提问来源于stack exchange,提问作者Tarek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:21:12