You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从多页面抓取汽车故障码并生成指定表格

多页面故障码抓取与表格生成方案

核心思路

  1. 批量管理目标页面URL,循环遍历每个页面执行抓取逻辑
  2. 统一收集每个页面的结构化数据(h2作为键,对应p/ul文本作为值)
  3. 利用Pandas将收集到的所有数据转换为指定格式的表格

修正后的完整代码

from selenium import webdriver
from bs4 import BeautifulSoup
import pandas as pd

# 1. 定义所有要抓取的页面URL列表
page_urls = [
    "https://www.somelink.com/page1",
    "https://www.somelink.com/page2",
    "https://www.somelink.com/page3",
    "https://www.somelink.com/page4"
]

# 初始化浏览器驱动(根据实际使用的浏览器调整,比如Chrome/Firefox)
driver = webdriver.Chrome()
result_data = []

# 2. 循环抓取每个页面的数据
for url in page_urls:
    driver.get(url)
    soup_subpage = BeautifulSoup(driver.page_source, "html.parser")
    # 抓取目标区域的h2和对应文本
    for content_block in soup_subpage.find_all('div', class_='col-md-8'):
        h2_titles = [h2.text.strip() for h2 in content_block.find_all('h2')]
        section_texts = [elem.text.strip() for elem in content_block.find_all(['p','ul'])]
        # 组装成字典,确保h2和文本一一对应
        page_data = dict(zip(h2_titles, section_texts))
        result_data.append(page_data)

# 关闭浏览器驱动
driver.quit()

# 3. 生成DataFrame并转换为指定格式的HTML表格
df = pd.DataFrame(result_data)
# 生成带指定类名的HTML表格
html_table = f'''
<div class="s-table-container">
<table class="s-table">
<thead>
<tr>
{''.join([f'<th>{col}</th>' for col in df.columns])}
</tr>
</thead>
<tbody>
{df.apply(lambda row: f'<tr>{''.join([f"<td>{cell}</td>" for cell in row])}</tr>', axis=1).str.cat(sep='\n')}
</tbody>
</table>
</div>
'''

# 可以将HTML保存到文件或直接使用
print(html_table)

关键说明

  • URL批量管理:把所有需要抓取的页面URL放到page_urls列表中,新增页面只需添加URL即可
  • 数据一致性:确保每个页面的h2标签数量和对应p/ul文本数量一致,否则zip会截断到较短的长度
  • 语法修正:原代码中pd.DataFrame([data]. columns = data.keys()}存在语法错误,已修正为标准的DataFrame初始化方式
  • HTML样式匹配:生成的HTML表格完全匹配你需要的class属性结构,可直接嵌入页面使用

内容的提问来源于stack exchange,提问作者MI6

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 06:20:45