You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python BeautifulSoup爬取表格存DataFrame数据缺失问题

爬取目标
现有实现代码
# Scraping a Data Table from the CHD Website
import requests
from bs4 import BeautifulSoup
import pandas as pd

# Load CHD Website HTML code
result = requests.get(current_url, verify=False, headers={'User-Agent' : "Magic Browser"})

# Check and see if the page successfully loaded
result_status = result.status_code
                      
if result.status_code == 200:
    # Extract the HTML code and pass it through beautiful soup
    source = result.content
    document = BeautifulSoup(source, 'lxml')

    # Since each page has one table for each product, we can use the table attribute to find the table
    check = 0
    table = document.find("table")
    
    while check <= 0:
        # Check to make sure that you got the right table by checking whether the text within the first header title is 'INGREDIENT'
        if table.find("span").get_text() == "INGREDIENT NAME":
            check += 1
        else:
            table = table.find_next("table")
            
    # Since HTML uses tr for rows, we can use find all to get our rows
    rows = table.find_all('span', style ='font-size:13px;font-family:"Arial",sans-serif;')
        
    cells_names = []
    # Loop through the rows
    for row in rows[3:]:
        bar = row.find('span', style ='font-size:13px;font-family:"Arial",sans-serif;')
        bar_text = row.get_text(strip = True)
        cells_names.append(bar_text)
        
    
    data_pandas = pd.DataFrame(cells_names, columns = ['Ingredients'])
    # return data_pandas

else:
    # Print out an error if the result status is not 200
    print("Status error" + "  " + str(result_status) + "has occurred!")
遇到的问题

运行代码后生成的DataFrame缺失lubricant/emulsifer相关内容,排查确认原因是:对应内容所在的<span>标签style属性额外增加了color:black;background:white字段,现有代码采用精确匹配完整style属性值的逻辑,无法选中这部分元素,导致数据遗漏。

解决方案

问题核心是精确匹配style属性的逻辑容错性极差,只要前端给标签加样式、调整style内属性顺序、增减空格,就会匹配失败。两种可直接落地的修复方式:

  • 方案1:替换精确style匹配为模糊匹配
    不用要求style属性完全等于固定字符串,只要span的style包含目标字体大小、字体名两个核心特征就选中元素,用lambda函数做自定义匹配即可,替换原来查找rows的代码:

    rows = table.find_all(
        'span', 
        style=lambda style_val: style_val and 'font-size:13px' in style_val and 'Arial' in style_val
    )
    

    注意:循环内bar = row.find('span', style=...)是冗余代码,外层已经定位到符合要求的span,不需要再向内查找子span,直接取当前row的文本即可。

  • 方案2:完全放弃依赖style属性匹配(稳定性更高)
    style是前端展示层属性,随时可能调整,爬取结构化数据优先依赖标签的层级结构。已经定位到正确的成分表table后,直接遍历表格的所有行<tr>提取文本即可,完全不用管内部span加了什么样式:

    # 定位到正确table后,直接取所有表格行
    all_tr = table.find_all('tr')
    cells_names = []
    # 跳过前3行表头
    for tr in all_tr[3:]:
        ingredient_text = tr.get_text(strip=True)
        # 过滤空行
        if ingredient_text:
            cells_names.append(ingredient_text)
    
    data_pandas = pd.DataFrame(cells_names, columns = ['Ingredients'])
    

    该写法不受内部标签样式、标签结构微调的影响,长期运行漏数据的概率远低于匹配style的方案。


内容的提问来源于stack exchange,提问作者bolbolnuggets

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 13:18:22