You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup提取Shopee商品URL中的特定数字串

从Shopee商品URL提取「数字.数字」格式字符串的正确方法

为什么你的代码失败

你用area.find('a', text='-i.')匹配不到目标标签,是因为这个写法是找**文本内容完全等于"-i."**的<a>标签,但Shopee商品链接的-i.是在<a>的href属性里,不是标签的文本内容,自然匹配失败。

两种可行实现方式

方式1:正则匹配(推荐,适配URL格式变化)

先提取<a>标签的href属性,再用正则精准匹配目标字符串:

import re
from bs4 import BeautifulSoup

# 假设你已通过爬虫获取到目标页面的HTML
soup = BeautifulSoup(your_html_content, 'html.parser')

# 先定位到包含商品链接的容器(替换成你实际使用的选择器,比如商品卡片的class)
item_container = soup.find('div', class_='shopee-search-item-result__item')
if item_container:
    # 获取商品链接的a标签
    item_link = item_container.find('a')
    if item_link and 'href' in item_link.attrs:
        url = item_link['href']
        # 匹配URL中"i.数字.数字"的部分,捕获数字串
        match_result = re.search(r'i\.(\d+\.\d+)', url)
        if match_result:
            target_str = match_result.group(1)
            print(target_str)  # 输出:611674069.14413534248

方式2:字符串分割(适合格式稳定的场景)

如果Shopee商品URL的格式长期稳定(始终以-i.数字.数字?作为参数前缀),可以直接用字符串分割提取:

# 接上面的代码,获取到item_link的href后
url = item_link['href']
# 先按"-i."拆分取后半段,再按"?"拆分取前半段
target_str = url.split('-i.')[-1].split('?')[0]
print(target_str)

关键注意点

  • 定位<a>标签时,要基于稳定的属性(比如class、data属性),不要依赖标签文本或不稳定的层级;
  • 如果页面是JS动态渲染的(Shopee多数商品列表是动态加载),Beautiful Soup无法直接抓取到渲染后的链接,这时需要搭配Selenium、Playwright等工具处理动态内容。

内容的提问来源于stack exchange,提问作者Rinukz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 10:40:48