如何用BeautifulSoup高效解析HTML获取无冗余前缀的商品标题?
解决eBay商品标题解析中"Details about "前缀问题的高效方法
这是个很常见的eBay页面解析问题,我来给你几个不需要手动截取文本的高效解决方案,核心思路是精准定位到真正的商品标题文本节点,而不是直接取整个h1标签的文本内容。
问题根源
你遇到的情况是因为eBay把"Details about "放在了一个隐藏的<span class="g-hdn">标签里,这个标签和商品标题的文本节点都是<h1 class="it-ttl">的子节点:
item.string返回None是因为<h1>包含多个子节点(span元素+文本节点),只有当标签仅包含单个文本子节点时,.string才会返回有效内容;item.text会把所有子节点的文本拼接在一起,所以就带上了前缀。
解决方案1:直接取h1标签的最后一个子节点
根据页面结构,商品标题文本是<h1>的最后一个子节点,直接提取它即可:
import requests from bs4 import BeautifulSoup url1 = "https://www.ebay.com/itm/Big-Boss-Air-Fryer-Healthy-1300-Watt-Super-Sized-16-Quart-Fryer-5-Colors-NEW/122454150244? epid=2254405949&hash=item1c82d60c64:m:mqfT2XbgveSevmN5MV1iysg" def get_single_item_data(item_url): source_code = requests.get(item_url) plain_text = source_code.text soup = BeautifulSoup(plain_text, 'html.parser') title_tag = soup.find('h1', {'class':'it-ttl'}) # 提取h1标签的最后一个子节点(即商品标题文本)并去除首尾空白 product_title = title_tag.contents[-1].strip() print(product_title) get_single_item_data(url1)
解决方案2:利用stripped_strings过滤前缀
stripped_strings会返回标签下所有非空白的文本片段,我们只需要跳过第一个片段(就是"Details about")即可:
import requests from bs4 import BeautifulSoup url1 = "https://www.ebay.com/itm/Big-Boss-Air-Fryer-Healthy-1300-Watt-Super-Sized-16-Quart-Fryer-5-Colors-NEW/122454150244? epid=2254405949&hash=item1c82d60c64:m:mqfT2XbgveSevmN5MV1iysg" def get_single_item_data(item_url): source_code = requests.get(item_url) plain_text = source_code.text soup = BeautifulSoup(plain_text, 'html.parser') title_tag = soup.find('h1', {'class':'it-ttl'}) # 转换为列表后跳过第一个元素,再拼接成完整标题 title_fragments = list(title_tag.stripped_strings) product_title = ' '.join(title_fragments[1:]) print(product_title) get_single_item_data(url1)
解决方案3:定位隐藏span的下一个兄弟节点
直接找到隐藏的<span class="g-hdn">,然后取它的next_sibling(也就是紧跟在后面的商品标题文本):
import requests from bs4 import BeautifulSoup url1 = "https://www.ebay.com/itm/Big-Boss-Air-Fryer-Healthy-1300-Watt-Super-Sized-16-Quart-Fryer-5-Colors-NEW/122454150244? epid=2254405949&hash=item1c82d60c64:m:mqfT2XbgveSevmN5MV1iysg" def get_single_item_data(item_url): source_code = requests.get(item_url) plain_text = source_code.text soup = BeautifulSoup(plain_text, 'html.parser') hidden_prefix_span = soup.find('span', {'class':'g-hdn'}) # 取span的下一个兄弟节点(文本节点)并去除空白 product_title = hidden_prefix_span.next_sibling.strip() print(product_title) get_single_item_data(url1)
这三种方法都不需要手动截取字符串,而是通过DOM节点定位直接获取目标内容,比文本截取更稳定,也更符合网页解析的最佳实践。
内容的提问来源于stack exchange,提问作者Joe Chan
相关产品推荐
相关产品推荐

