You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改BeautifulSoup代码仅获取当前行内的<th>标签而非全部<th>标签

问题根源

你代码中使用的find_all_next()方法的作用是查找当前标签之后、整个文档内所有匹配的元素,而非当前标签的子元素,所以才会返回所有表格的<th>标签。

修改方案

把find_all_next()替换为find_all()即可,find_all()默认只会查找当前节点的所有后代匹配元素,刚好符合你要找当前表格内的行、当前行内的<th>的需求。

同时针对你需要每个表格单独生成列表的需求,可以参考下面的完整代码:

def parse_html(self):
    """ Parse the html file """
    # 建议加上encoding避免中文乱码
    with open(self.html_path, encoding='utf-8') as f:
        soup = BeautifulSoup(f, 'html.parser')

    tables = soup.find_all('table')
    # 存储所有表格的结果,每个子列表对应一个表格的内容
    all_tables_data = []
    
    for table in tables:
        current_table = []
        # 只查找当前表格下的所有tr
        rows = table.find_all('tr')
        for row in rows:
            # 只查找当前行下的所有th
            cols = row.find_all('th')
            # 有th的行就是表头行
            if cols:
                header_row = []
                for col in cols:
                    # 提取单元格文本,自动合并多个p标签的内容,用空格分隔
                    cell_text = col.get_text(strip=True, separator=' ')
                    header_row.append(cell_text)
                current_table.append(header_row)
            # 需要处理普通数据行可以在这里补充逻辑
            # else:
            #     tds = row.find_all('td')
            #     后续处理逻辑
        all_tables_data.append(current_table)
    
    # 可输出验证结果
    for idx, table in enumerate(all_tables_data):
        print(f"第{idx+1}个表格的表头:", table)
    
    return all_tables_data

补充优化

如果只需要提取每个表格带class="header"的表头行,还可以进一步简化查找逻辑,不用遍历所有行:

for table in tables:
    header_row = table.select_one('tr.header')
    if header_row:
        cols = header_row.find_all('th')
        # 后续处理逻辑和上面一致

内容的提问来源于stack exchange,提问作者user16549853

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 13:18:00