You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup高效解析HTML获取无冗余前缀的商品标题?

解决eBay商品标题解析中"Details about "前缀问题的高效方法

这是个很常见的eBay页面解析问题,我来给你几个不需要手动截取文本的高效解决方案,核心思路是精准定位到真正的商品标题文本节点,而不是直接取整个h1标签的文本内容。

问题根源

你遇到的情况是因为eBay把"Details about "放在了一个隐藏的<span class="g-hdn">标签里,这个标签和商品标题的文本节点都是<h1 class="it-ttl">的子节点:

  • item.string返回None是因为<h1>包含多个子节点(span元素+文本节点),只有当标签仅包含单个文本子节点时,.string才会返回有效内容;
  • item.text会把所有子节点的文本拼接在一起,所以就带上了前缀。

解决方案1:直接取h1标签的最后一个子节点

根据页面结构,商品标题文本是<h1>的最后一个子节点,直接提取它即可:

import requests
from bs4 import BeautifulSoup

url1 = "https://www.ebay.com/itm/Big-Boss-Air-Fryer-Healthy-1300-Watt-Super-Sized-16-Quart-Fryer-5-Colors-NEW/122454150244? epid=2254405949&hash=item1c82d60c64:m:mqfT2XbgveSevmN5MV1iysg"

def get_single_item_data(item_url):
    source_code = requests.get(item_url)
    plain_text = source_code.text
    soup = BeautifulSoup(plain_text, 'html.parser')
    title_tag = soup.find('h1', {'class':'it-ttl'})
    # 提取h1标签的最后一个子节点(即商品标题文本)并去除首尾空白
    product_title = title_tag.contents[-1].strip()
    print(product_title)

get_single_item_data(url1)

解决方案2:利用stripped_strings过滤前缀

stripped_strings会返回标签下所有非空白的文本片段,我们只需要跳过第一个片段(就是"Details about")即可:

import requests
from bs4 import BeautifulSoup

url1 = "https://www.ebay.com/itm/Big-Boss-Air-Fryer-Healthy-1300-Watt-Super-Sized-16-Quart-Fryer-5-Colors-NEW/122454150244? epid=2254405949&hash=item1c82d60c64:m:mqfT2XbgveSevmN5MV1iysg"

def get_single_item_data(item_url):
    source_code = requests.get(item_url)
    plain_text = source_code.text
    soup = BeautifulSoup(plain_text, 'html.parser')
    title_tag = soup.find('h1', {'class':'it-ttl'})
    # 转换为列表后跳过第一个元素,再拼接成完整标题
    title_fragments = list(title_tag.stripped_strings)
    product_title = ' '.join(title_fragments[1:])
    print(product_title)

get_single_item_data(url1)

解决方案3:定位隐藏span的下一个兄弟节点

直接找到隐藏的<span class="g-hdn">,然后取它的next_sibling(也就是紧跟在后面的商品标题文本):

import requests
from bs4 import BeautifulSoup

url1 = "https://www.ebay.com/itm/Big-Boss-Air-Fryer-Healthy-1300-Watt-Super-Sized-16-Quart-Fryer-5-Colors-NEW/122454150244? epid=2254405949&hash=item1c82d60c64:m:mqfT2XbgveSevmN5MV1iysg"

def get_single_item_data(item_url):
    source_code = requests.get(item_url)
    plain_text = source_code.text
    soup = BeautifulSoup(plain_text, 'html.parser')
    hidden_prefix_span = soup.find('span', {'class':'g-hdn'})
    # 取span的下一个兄弟节点(文本节点)并去除空白
    product_title = hidden_prefix_span.next_sibling.strip()
    print(product_title)

get_single_item_data(url1)

这三种方法都不需要手动截取字符串,而是通过DOM节点定位直接获取目标内容,比文本截取更稳定,也更符合网页解析的最佳实践。

内容的提问来源于stack exchange,提问作者Joe Chan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:01:43