You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup提取页面链接、标题与颜色文本的技术问题

解决Beautiful Soup提取网页链接、标题与颜色文本的问题

问题背景

刚上手Beautiful Soup的时候,我在网页数据提取上碰了壁:现有的代码能找到turbolink_scroller类div下的所有article元素,但输出的是带完整HTML标签的<h1>和<p>内容。我需要单独把每个商品的链接、<h1>里的标题文本、<p>里的颜色文本提取出来,分别存入link、name、color变量中。

目标页面结构示例

<div class="turbolink_scroller" id="container" style="opacity: 1;"> 
  <article> 
    <div class="inner-article"> 
      <a style="height:150px;" href="/shop/jackets/h21snm5ld/jick90fel"> 
        <img width="150" height="150" src="//assets.supremenewyork.com/146917/vi/MCHFhUqvN0w.jpg" alt="Mchfhuqvn0w"> 
        <div class="sold_out_tag" style="">sold out</div> 
      </a> 
      <h1><a class="name-link" href="/shop/jackets/h21snm5ld/jick90fel">NY Tapestry Denim Chore Coat</a></h1> 
      <p><a class="name-link" href="/shop/jackets/h21snm5ld/jick90fel">Maroon</a></p> 
    </div> 
  </article> 
  <article></article> 
  <article></article> 
  <article></article> 
</div>

原代码及问题输出

原代码:

article_name_list = soup.find(class_='turbolink_scroller') #find all links in the div
article_name_list_items = article_name_list.find_all('article') #loop to print all out
for article_name in article_name_list_items:
    names = article_name.find('h1')
    color = article_name.find('p')
    print(names)
    print(color)

输出结果(包含冗余HTML标签):

<h1><a class="name-link" href="/shop/jackets/gw1diqgyr/km21a8hnc">Gonz Logo Coaches Jacket </a></h1>
<p><a class="name-link" href="/shop/jackets/gw1diqgyr/km21a8hnc">Red</a></p>

解决方案代码

要拿到纯净的文本和链接,需要深入嵌套的<a>标签,用get_text()提取文本内容,用get('href')获取链接属性:

article_name_list = soup.find(class_='turbolink_scroller') # 定位目标div容器
article_name_list_items = article_name_list.find_all('article') # 获取所有article元素

# 遍历每个article提取对应数据
for article_name in article_name_list_items:
    # 提取商品链接
    link = article_name.find('h1').find('a').get('href')
    # 提取标题纯文本
    name = article_name.find('h1').find('a').get_text()
    # 提取颜色纯文本
    color = article_name.find('p').find('a').get_text()
    
    # 打印或按需使用变量
    print(f"标题: {name}")
    print(f"颜色: {color}")
    print(f"链接: {link}")
    print("---")

关键说明

  • find('h1').find('a'):逐层定位到承载内容的<a>标签,因为标题和颜色文本都嵌套在这个标签内
  • get('href'):直接获取<a>标签的href属性值,也就是商品的跳转链接
  • get_text():自动剥离HTML标签,提取标签内的纯文本内容

内容的提问来源于stack exchange,提问作者lucky simon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:28:15