You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup获取表格标题并追加到每行数据中

解决方案

修改函数逻辑

原函数仅提取表格行数据,现在需要新增页面title获取逻辑,并将标题与每行文本内容合并。同时补充异常处理,避免因页面结构异常导致报错。

修改后的函数

from urllib.request import urlopen
from bs4 import BeautifulSoup

def get_team_table_with_title(url):
    page = urlopen(url)
    soup = BeautifulSoup(page, 'lxml')
    
    # 获取页面标题,处理无标题的情况
    page_title = soup.title.string.strip() if soup.title else "Unknown Title"
    
    # 定位目标表格,先判断表格是否存在
    target_table = soup.find("table", class_="datatable")
    if not target_table:
        return []
    
    # 遍历行并拼接标题与行文本
    formatted_rows = []
    for row in target_table.find_all("tr"):
        # 提取该行所有单元格的文本,去除多余空格并合并
        row_content = " ".join([cell.get_text(strip=True) for cell in row.find_all(["td", "th"])])
        # 跳过空行,拼接标题与行内容
        if row_content:
            formatted_rows.append(f"{page_title} | {row_content}")
    
    return formatted_rows

遍历URL列表收集最终数据

遍历links_all中的每个URL,调用修改后的函数,将结果汇总到table_data:

# 替换为你的实际URL列表
links_all = ["https://example.com/team1", "https://example.com/team2"]
table_data = []

for url in links_all:
    rows_with_title = get_team_table_with_title(url)
    table_data.extend(rows_with_title)

关键说明

  • 新增title空值处理,避免页面无标题时抛出异常
  • 提取行文本时同时处理表头<th>和数据<td>,确保内容完整
  • 加入表格存在性判断,防止找不到目标表格导致程序崩溃
  • 最终table_data的每个元素格式为:"页面标题 | 行内容",符合示例要求

内容的提问来源于stack exchange,提问作者ElphiusMostafa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 12:17:19