网页爬取中的多级标签存在性检查——提升Python代码可读性
嘿,针对你开发Python爬虫的需求——爬取同模板商品页面、获取全量数据、实现多级标签检查还得提升代码可读性,我给你整了一套实用的方案,直接就能上手改:
核心思路拆解
首先得解决两个核心问题:
- 多级标签存在性检查:页面难免会有缺失的元素(比如某个商品没填制造商),直接链式调用
find会触发AttributeError,所以必须逐层级确认节点存在后再提取内容。 - 代码可读性提升:把不同字段的提取逻辑拆成独立函数,用见名知意的命名,加上必要注释,避免一坨代码堆在一起。
完整代码实现
我用requests做网络请求,BeautifulSoup做页面解析,代码如下:
import requests from bs4 import BeautifulSoup from typing import Dict, Optional def get_product_name(soup: BeautifulSoup) -> Optional[str]: """提取商品名称:从content div下的h1标签获取""" content_div = soup.find('div', id='content') if not content_div: return None product_title = content_div.find('h1') return product_title.get_text(strip=True) if product_title else None def get_manufacturer(soup: BeautifulSoup) -> Optional[str]: """提取制造商:从properties表格的manufacturer-row行获取""" properties_table = soup.find('table', id='properties') if not properties_table: return None manufacturer_row = properties_table.find('tr', id='manufacturer-row') if not manufacturer_row: return None manufacturer_value = manufacturer_row.find('td') return manufacturer_value.get_text(strip=True) if manufacturer_value else None def get_product_price(soup: BeautifulSoup) -> Optional[str]: """提取商品价格:这里假设价格在properties表格的price-row行,你可以根据实际结构调整""" properties_table = soup.find('table', id='properties') if not properties_table: return None price_row = properties_table.find('tr', id='price-row') if not price_row: return None price_value = price_row.find('td') return price_value.get_text(strip=True) if price_value else None def get_product_description(soup: BeautifulSoup) -> Optional[str]: """提取商品描述:假设描述在content div下的description类标签,根据实际结构调整""" content_div = soup.find('div', id='content') if not content_div: return None description_tag = content_div.find('div', class_='description') return description_tag.get_text(strip=True) if description_tag else None def scrape_single_page(url: str) -> Dict[str, Optional[str]]: """爬取单个商品页面,返回结构化的商品数据""" try: response = requests.get(url, timeout=10) response.raise_for_status() # 触发HTTP错误异常 soup = BeautifulSoup(response.text, 'html.parser') return { 'product_name': get_product_name(soup), 'manufacturer': get_manufacturer(soup), 'price': get_product_price(soup), 'description': get_product_description(soup) } except Exception as e: print(f"爬取页面 {url} 失败:{str(e)}") return { 'product_name': None, 'manufacturer': None, 'price': None, 'description': None } def scrape_multiple_pages(urls: list[str]) -> list[Dict[str, Optional[str]]]: """批量爬取多个商品页面""" all_product_data = [] for url in urls: product_data = scrape_single_page(url) all_product_data.append(product_data) return all_product_data # 示例使用 if __name__ == '__main__': # 替换成你的目标页面URL列表 target_urls = [ 'https://example.com/product-page-1', 'https://example.com/product-page-2', 'https://example.com/product-page-3' ] products = scrape_multiple_pages(target_urls) for idx, product in enumerate(products, 1): print(f"第{idx}个商品数据:") print(product) print("---")
关键优势说明
- 多级检查防崩溃:每个提取函数都逐层级判断节点是否存在,比如找制造商时,先确认
properties表格存在,再找manufacturer-row行,最后找td标签,任何一级缺失都返回None,不会让爬虫直接报错终止。 - 模块化易维护:每个字段的提取逻辑单独封装,后续要修改价格的提取规则,只需要改
get_product_price函数就行,不用动其他代码。 - 可读性拉满:函数名(比如
get_product_name)和变量名都直观易懂,加上函数注释说明用途,新人接手也能快速看懂。 - 容错能力强:网络请求部分加了
try-except,处理超时、HTTP错误等异常,批量爬取时单个页面失败不会影响其他页面。
你可以根据实际页面的HTML结构,调整每个提取函数里的标签选择器(比如价格、描述的标签位置),适配你的目标网站。
内容的提问来源于stack exchange,提问作者stackrat
相关产品推荐
相关产品推荐

