You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取中的多级标签存在性检查——提升Python代码可读性

嘿,针对你开发Python爬虫的需求——爬取同模板商品页面、获取全量数据、实现多级标签检查还得提升代码可读性,我给你整了一套实用的方案,直接就能上手改:

核心思路拆解

首先得解决两个核心问题:

  1. 多级标签存在性检查:页面难免会有缺失的元素(比如某个商品没填制造商),直接链式调用find会触发AttributeError,所以必须逐层级确认节点存在后再提取内容。
  2. 代码可读性提升:把不同字段的提取逻辑拆成独立函数,用见名知意的命名,加上必要注释,避免一坨代码堆在一起。
完整代码实现

我用requests做网络请求,BeautifulSoup做页面解析,代码如下:

import requests
from bs4 import BeautifulSoup
from typing import Dict, Optional

def get_product_name(soup: BeautifulSoup) -> Optional[str]:
    """提取商品名称:从content div下的h1标签获取"""
    content_div = soup.find('div', id='content')
    if not content_div:
        return None
    product_title = content_div.find('h1')
    return product_title.get_text(strip=True) if product_title else None

def get_manufacturer(soup: BeautifulSoup) -> Optional[str]:
    """提取制造商:从properties表格的manufacturer-row行获取"""
    properties_table = soup.find('table', id='properties')
    if not properties_table:
        return None
    manufacturer_row = properties_table.find('tr', id='manufacturer-row')
    if not manufacturer_row:
        return None
    manufacturer_value = manufacturer_row.find('td')
    return manufacturer_value.get_text(strip=True) if manufacturer_value else None

def get_product_price(soup: BeautifulSoup) -> Optional[str]:
    """提取商品价格:这里假设价格在properties表格的price-row行,你可以根据实际结构调整"""
    properties_table = soup.find('table', id='properties')
    if not properties_table:
        return None
    price_row = properties_table.find('tr', id='price-row')
    if not price_row:
        return None
    price_value = price_row.find('td')
    return price_value.get_text(strip=True) if price_value else None

def get_product_description(soup: BeautifulSoup) -> Optional[str]:
    """提取商品描述:假设描述在content div下的description类标签,根据实际结构调整"""
    content_div = soup.find('div', id='content')
    if not content_div:
        return None
    description_tag = content_div.find('div', class_='description')
    return description_tag.get_text(strip=True) if description_tag else None

def scrape_single_page(url: str) -> Dict[str, Optional[str]]:
    """爬取单个商品页面,返回结构化的商品数据"""
    try:
        response = requests.get(url, timeout=10)
        response.raise_for_status()  # 触发HTTP错误异常
        soup = BeautifulSoup(response.text, 'html.parser')
        
        return {
            'product_name': get_product_name(soup),
            'manufacturer': get_manufacturer(soup),
            'price': get_product_price(soup),
            'description': get_product_description(soup)
        }
    except Exception as e:
        print(f"爬取页面 {url} 失败:{str(e)}")
        return {
            'product_name': None,
            'manufacturer': None,
            'price': None,
            'description': None
        }

def scrape_multiple_pages(urls: list[str]) -> list[Dict[str, Optional[str]]]:
    """批量爬取多个商品页面"""
    all_product_data = []
    for url in urls:
        product_data = scrape_single_page(url)
        all_product_data.append(product_data)
    return all_product_data

# 示例使用
if __name__ == '__main__':
    # 替换成你的目标页面URL列表
    target_urls = [
        'https://example.com/product-page-1',
        'https://example.com/product-page-2',
        'https://example.com/product-page-3'
    ]
    
    products = scrape_multiple_pages(target_urls)
    for idx, product in enumerate(products, 1):
        print(f"第{idx}个商品数据:")
        print(product)
        print("---")
关键优势说明
  • 多级检查防崩溃:每个提取函数都逐层级判断节点是否存在,比如找制造商时,先确认properties表格存在,再找manufacturer-row行,最后找td标签,任何一级缺失都返回None,不会让爬虫直接报错终止。
  • 模块化易维护:每个字段的提取逻辑单独封装,后续要修改价格的提取规则,只需要改get_product_price函数就行,不用动其他代码。
  • 可读性拉满:函数名(比如get_product_name)和变量名都直观易懂,加上函数注释说明用途,新人接手也能快速看懂。
  • 容错能力强:网络请求部分加了try-except,处理超时、HTTP错误等异常,批量爬取时单个页面失败不会影响其他页面。

你可以根据实际页面的HTML结构,调整每个提取函数里的标签选择器(比如价格、描述的标签位置),适配你的目标网站。

内容的提问来源于stack exchange,提问作者stackrat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:24:25