如何解决Python BeautifulSoup爬取表格存DataFrame数据缺失问题
爬取目标
- 待爬取页面为Church & Dwight品牌的止汗露成分披露页:https://churchdwight.com/ingredient-disclosure/antiperspirant-deodorant/40002569-ultramax-clear-gel-cool-blast.aspx
现有实现代码
# Scraping a Data Table from the CHD Website import requests from bs4 import BeautifulSoup import pandas as pd # Load CHD Website HTML code result = requests.get(current_url, verify=False, headers={'User-Agent' : "Magic Browser"}) # Check and see if the page successfully loaded result_status = result.status_code if result.status_code == 200: # Extract the HTML code and pass it through beautiful soup source = result.content document = BeautifulSoup(source, 'lxml') # Since each page has one table for each product, we can use the table attribute to find the table check = 0 table = document.find("table") while check <= 0: # Check to make sure that you got the right table by checking whether the text within the first header title is 'INGREDIENT' if table.find("span").get_text() == "INGREDIENT NAME": check += 1 else: table = table.find_next("table") # Since HTML uses tr for rows, we can use find all to get our rows rows = table.find_all('span', style ='font-size:13px;font-family:"Arial",sans-serif;') cells_names = [] # Loop through the rows for row in rows[3:]: bar = row.find('span', style ='font-size:13px;font-family:"Arial",sans-serif;') bar_text = row.get_text(strip = True) cells_names.append(bar_text) data_pandas = pd.DataFrame(cells_names, columns = ['Ingredients']) # return data_pandas else: # Print out an error if the result status is not 200 print("Status error" + " " + str(result_status) + "has occurred!")
遇到的问题
运行代码后生成的DataFrame缺失lubricant/emulsifer相关内容,排查确认原因是:对应内容所在的<span>标签style属性额外增加了color:black;background:white字段,现有代码采用精确匹配完整style属性值的逻辑,无法选中这部分元素,导致数据遗漏。
解决方案
问题核心是精确匹配style属性的逻辑容错性极差,只要前端给标签加样式、调整style内属性顺序、增减空格,就会匹配失败。两种可直接落地的修复方式:
方案1:替换精确style匹配为模糊匹配
不用要求style属性完全等于固定字符串,只要span的style包含目标字体大小、字体名两个核心特征就选中元素,用lambda函数做自定义匹配即可,替换原来查找rows的代码:rows = table.find_all( 'span', style=lambda style_val: style_val and 'font-size:13px' in style_val and 'Arial' in style_val )注意:循环内
bar = row.find('span', style=...)是冗余代码,外层已经定位到符合要求的span,不需要再向内查找子span,直接取当前row的文本即可。方案2:完全放弃依赖style属性匹配(稳定性更高)
style是前端展示层属性,随时可能调整,爬取结构化数据优先依赖标签的层级结构。已经定位到正确的成分表table后,直接遍历表格的所有行<tr>提取文本即可,完全不用管内部span加了什么样式:# 定位到正确table后,直接取所有表格行 all_tr = table.find_all('tr') cells_names = [] # 跳过前3行表头 for tr in all_tr[3:]: ingredient_text = tr.get_text(strip=True) # 过滤空行 if ingredient_text: cells_names.append(ingredient_text) data_pandas = pd.DataFrame(cells_names, columns = ['Ingredients'])该写法不受内部标签样式、标签结构微调的影响,长期运行漏数据的概率远低于匹配style的方案。
内容的提问来源于stack exchange,提问作者bolbolnuggets
相关产品推荐
相关产品推荐

