如何用BeautifulSoup通过summary和width抓取无标识表格?XPath兼容吗?
问题描述
我尝试从FDA2015年召回档案的归档页面抓取表格,目标表格没有id或class属性,仅带有width和summary属性。想请教有没有办法抓取这个表格?能不能用XPath?另外听说XPath和BeautifulSoup不兼容,希望这个说法是错的。
目标表格的代码片段:
<table width="100%" cellpadding="3" border="1" summary="Layout showing RecallTest table with 6 columns: Date,Brand Name,Product Description,Reason/Problem,Company,Details/Photo" style="margin-bottom:28px"> <thead> <tr> <th scope="col" data-type="numeric" data-toggle="true"> Date </th> </tr> </thead> <tbody>
我目前的代码:
import requests from bs4 import BeautifulSoup link = 'http://wayback.archive-it.org/7993/20170110233205/http://www.fda.gov/Safety/Recalls/ArchiveRecalls/2015/default.htm' page = 15 pdf = [] for p in range(1,page+1): l = link + '?page='+str(p) # Downloading contents of the web page data = requests.get(l).text # Creating BeautifulSoup object soup = BeautifulSoup(data, 'html.parser') tables = soup.find_all('table') table = soup.find('table', INSERT XPATH EXPRESSION) df = pd.DataFrame(columns = ['date','brand','descr','reason','company']) for row in table.tbody.find_all('tr'): # Find all data for each column columns = row.find_all('td') if columns != []: date = columns[0].text.strip()
解决方案
1. 直接用BeautifulSoup通过属性定位表格,无需XPath
BeautifulSoup原生确实不支持XPath,但完全没必要用XPath——你可以直接通过表格的summary属性精准定位,这个属性的内容是唯一的,比width更可靠:
table = soup.find('table', summary="Layout showing RecallTest table with 6 columns: Date,Brand Name,Product Description,Reason/Problem,Company,Details/Photo")
如果担心summary内容有细微差异,也可以结合width和style属性双重过滤:
table = soup.find('table', {'width': '100%', 'style': 'margin-bottom:28px'})
2. 关于XPath和BeautifulSoup的兼容性
没错,BeautifulSoup本身不支持XPath语法。如果非要用XPath,你需要把BeautifulSoup的对象转成lxml节点再查询,但这属于多此一举——用原生的BeautifulSoup方法已经能完美解决你的问题。
3. 修正后的完整可运行代码
我帮你完善了代码逻辑,修复了缩进问题,添加了异常处理,优化了数据收集方式:
import requests import pandas as pd from bs4 import BeautifulSoup base_link = 'http://wayback.archive-it.org/7993/20170110233205/http://www.fda.gov/Safety/Recalls/ArchiveRecalls/2015/default.htm' total_pages = 15 all_records = [] for page_num in range(1, total_pages + 1): current_url = f"{base_link}?page={page_num}" # 处理请求异常,避免单个页面失败导致脚本中断 try: resp = requests.get(current_url, timeout=10) resp.raise_for_status() html_content = resp.text except requests.exceptions.RequestException as err: print(f"第{page_num}页请求失败: {err}") continue soup = BeautifulSoup(html_content, 'html.parser') # 通过summary属性定位目标表格 target_table = soup.find('table', summary="Layout showing RecallTest table with 6 columns: Date,Brand Name,Product Description,Reason/Problem,Company,Details/Photo") if not target_table: print(f"第{page_num}页未找到目标表格") continue # 遍历表格行提取数据 for row in target_table.tbody.find_all('tr'): cells = row.find_all('td') # 确保有足够的列数再提取 if len(cells) >= 5: record = { 'date': cells[0].text.strip(), 'brand': cells[1].text.strip(), 'descr': cells[2].text.strip(), 'reason': cells[3].text.strip(), 'company': cells[4].text.strip() } all_records.append(record) # 转换为DataFrame并输出 df = pd.DataFrame(all_records, columns=['date','brand','descr','reason','company']) print(df.head()) # 可选:保存到CSV文件 # df.to_csv('fda_2015_recalls.csv', index=False)
重要提示
- 优先用
summary属性定位,因为它的描述是唯一的,不会和页面其他表格混淆 - 不要在循环里反复创建DataFrame,先把所有数据存入列表再一次性转换,性能会好很多
- 添加异常处理能让脚本更健壮,避免因网络问题或页面结构变化直接崩溃
内容的提问来源于stack exchange,提问作者rokman54
相关产品推荐
相关产品推荐

