如何使用BeautifulSoup抓取网页中Product Dimensions字段信息
用BeautifulSoup提取Product Dimensions字段实现方案
以下是针对电商类页面规格字段的通用抓取代码,适配绝大多数常见的页面结构:
核心实现逻辑
from bs4 import BeautifulSoup import re # 此处soup为你已初始化完成的BeautifulSoup对象 product_info = {} # 适配列表/自由布局类结构 for text_node in soup.find_all(string=True): key = text_node.strip() if key == "Product Dimensions": # 提取相邻下一个有效节点的文本,清洗多余空格换行 dim_value = text_node.find_next().get_text(strip=True) dim_value = re.sub(r'\s+', ' ', dim_value) product_info["Product Dimensions"] = dim_value break # 若上面逻辑未匹配到,补充适配表格结构 if not product_info: for tr in soup.find_all("tr"): cells = tr.find_all("td") if len(cells) < 2: continue key = cells[0].get_text(strip=True) if key == "Product Dimensions": dim_value = cells[1].get_text(strip=True) dim_value = re.sub(r'\s+', ' ', dim_value) product_info["Product Dimensions"] = dim_value break print(product_info)
输出结果
运行后即可得到你需要的结构化数据格式:{'Product Dimensions': '102.87 x 40.64 x 11.43 cm; 16.33 Kilograms'}
- 如果你的页面有特殊的标签嵌套规则,可以调整
find_next()的参数,指定具体标签名(比如find_next("span"))提升匹配精度
内容的提问来源于stack exchange,提问作者D_S
相关产品推荐
相关产品推荐

