You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup通过summary和width抓取无标识表格?XPath兼容吗?

问题描述

我尝试从FDA2015年召回档案的归档页面抓取表格,目标表格没有id或class属性,仅带有width和summary属性。想请教有没有办法抓取这个表格?能不能用XPath?另外听说XPath和BeautifulSoup不兼容,希望这个说法是错的。

目标表格的代码片段:

<table width="100%" cellpadding="3" border="1" summary="Layout showing RecallTest table with 6 columns: Date,Brand Name,Product Description,Reason/Problem,Company,Details/Photo" style="margin-bottom:28px">
          <thead>
            <tr>
                    <th scope="col" data-type="numeric" data-toggle="true"> Date </th>
            </tr>
          </thead>
          <tbody>

我目前的代码:

import requests
from bs4 import BeautifulSoup
link = 'http://wayback.archive-it.org/7993/20170110233205/http://www.fda.gov/Safety/Recalls/ArchiveRecalls/2015/default.htm'
page = 15
pdf = []
for p in range(1,page+1):
   l = link + '?page='+str(p)
    # Downloading contents of the web page
    data = requests.get(l).text
    # Creating BeautifulSoup object
    soup = BeautifulSoup(data, 'html.parser')
    tables = soup.find_all('table')
    table = soup.find('table', INSERT XPATH EXPRESSION)
    df = pd.DataFrame(columns = ['date','brand','descr','reason','company'])
    for row in table.tbody.find_all('tr'):    
        # Find all data for each column
        columns = row.find_all('td')
        if columns != []:
            date = columns[0].text.strip()

解决方案

1. 直接用BeautifulSoup通过属性定位表格,无需XPath

BeautifulSoup原生确实不支持XPath,但完全没必要用XPath——你可以直接通过表格的summary属性精准定位,这个属性的内容是唯一的,比width更可靠:

table = soup.find('table', summary="Layout showing RecallTest table with 6 columns: Date,Brand Name,Product Description,Reason/Problem,Company,Details/Photo")

如果担心summary内容有细微差异,也可以结合width和style属性双重过滤:

table = soup.find('table', {'width': '100%', 'style': 'margin-bottom:28px'})

2. 关于XPath和BeautifulSoup的兼容性

没错,BeautifulSoup本身不支持XPath语法。如果非要用XPath,你需要把BeautifulSoup的对象转成lxml节点再查询,但这属于多此一举——用原生的BeautifulSoup方法已经能完美解决你的问题。

3. 修正后的完整可运行代码

我帮你完善了代码逻辑,修复了缩进问题,添加了异常处理,优化了数据收集方式:

import requests
import pandas as pd
from bs4 import BeautifulSoup

base_link = 'http://wayback.archive-it.org/7993/20170110233205/http://www.fda.gov/Safety/Recalls/ArchiveRecalls/2015/default.htm'
total_pages = 15
all_records = []

for page_num in range(1, total_pages + 1):
    current_url = f"{base_link}?page={page_num}"
    # 处理请求异常,避免单个页面失败导致脚本中断
    try:
        resp = requests.get(current_url, timeout=10)
        resp.raise_for_status()
        html_content = resp.text
    except requests.exceptions.RequestException as err:
        print(f"第{page_num}页请求失败: {err}")
        continue
    
    soup = BeautifulSoup(html_content, 'html.parser')
    # 通过summary属性定位目标表格
    target_table = soup.find('table', summary="Layout showing RecallTest table with 6 columns: Date,Brand Name,Product Description,Reason/Problem,Company,Details/Photo")
    
    if not target_table:
        print(f"第{page_num}页未找到目标表格")
        continue
    
    # 遍历表格行提取数据
    for row in target_table.tbody.find_all('tr'):
        cells = row.find_all('td')
        # 确保有足够的列数再提取
        if len(cells) >= 5:
            record = {
                'date': cells[0].text.strip(),
                'brand': cells[1].text.strip(),
                'descr': cells[2].text.strip(),
                'reason': cells[3].text.strip(),
                'company': cells[4].text.strip()
            }
            all_records.append(record)

# 转换为DataFrame并输出
df = pd.DataFrame(all_records, columns=['date','brand','descr','reason','company'])
print(df.head())
# 可选:保存到CSV文件
# df.to_csv('fda_2015_recalls.csv', index=False)

重要提示

  • 优先用summary属性定位,因为它的描述是唯一的,不会和页面其他表格混淆
  • 不要在循环里反复创建DataFrame,先把所有数据存入列表再一次性转换,性能会好很多
  • 添加异常处理能让脚本更健壮,避免因网络问题或页面结构变化直接崩溃

内容的提问来源于stack exchange,提问作者rokman54

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 15:05:27