You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取多分类页面的URL参数配置问题求助

解决Jurong Point商店目录爬取的URL参数问题

问题背景

需要爬取Jurong Point商店目录中以下4个分类的内容:

  • Service
  • Food & Beverage
  • Fashion & Accessories
  • Electronics & Technology

原代码的URL模板无法正确填充参数,尤其Service分类的URL格式和其他分类差异明显,导致无法正确请求对应分类的页面。

问题分析

  1. URL参数不匹配:原代码使用固定模板https://www.jurongpoint.com.sg/store-directory/?level=&cate={}+%26+{},但只有带&的分类需要拆分,Service分类实际对应的参数是Services(而非原代码中的Service),不需要拆分。
  2. 分页逻辑混乱:原代码在循环页码时会被next_button修改全局的url变量,导致后续分类的请求出错。
  3. 内容解析冗余:原代码中嵌套的循环和重复查找DOM元素的逻辑可以优化。

修正后的代码

from bs4 import BeautifulSoup
import requests
from urllib.parse import quote

def parse():
    # 定义分类名称与对应URL参数的映射,修正Service为Services(网站实际参数)
    category_map = {
        "Services": "Services",
        "Food & Beverage": "Food & Beverage",
        "Fashion & Accessories": "Fashion & Accessories",
        "Electronics & Technology": "Electronics & Technology"
    }

    base_url = "https://www.jurongpoint.com.sg/store-directory/"

    for category_name, category_param in category_map.items():
        print(f"正在爬取分类: {category_name}")
        # 对分类参数进行URL编码
        encoded_cate = quote(category_param)
        current_url = f"{base_url}?level=&cate={encoded_cate}"
        
        while True:
            response = requests.get(current_url)
            soup = BeautifulSoup(response.text, "html.parser")
            
            # 解析商店名称和描述
            names = soup.find_all('tr', class_="clickable")
            shops = soup.find_all('div', class_="col-9")
            
            for n, k in zip(names, shops):
                try:
                    name = n.find_all('td')[1].text.strip().replace(' ', '')
                    desc = k.text.strip().replace(' ', '')
                    print(f"商店名称: {name}")
                    print(f"描述: {desc}\n")
                except IndexError as e:
                    print(f"解析出错: {e}")
            
            # 处理分页
            next_button = soup.select_one('.PagedList-skipToNext a')
            if next_button:
                # 拼接完整的下一页URL
                next_href = next_button.get('href')
                current_url = f"{base_url}{next_href}"
            else:
                print(f"{category_name} 分类爬取完成\n")
                break

parse()

代码说明

  1. 分类参数映射:直接定义分类名称与网站实际接受的参数映射,同时修正了Service为Services(网站实际使用的参数是Services)。
  2. URL编码:使用urllib.parse.quote对分类参数进行统一编码,避免手动拼接%26的麻烦,确保所有分类参数都符合URL规范。
  3. 分页逻辑优化:使用while True循环处理分页,每次请求后检查下一页按钮,拼接完整URL,避免全局变量污染。
  4. 解析逻辑简化:移除冗余的循环,直接查找所需的DOM元素,同时添加strip()处理文本前后空格,提升解析准确性。

内容的提问来源于stack exchange,提问作者Bee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 01:35:25