You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python与BS4:如何通过h2类获取对应a标签链接及文本

Solution for Your New Scraping Task

Got it, let's tweak your code to work with the new HTML structure you're dealing with. Since the <h2> has a clear class (top-sec-title) and it's nested inside the <a> tag you need the href from, we can approach this by first targeting the <h2> elements, then grabbing their parent <a> tags.

Here's the adjusted code:

from bs4 import BeautifulSoup

# Assume your HTML content is stored in a variable called html_content
soup = BeautifulSoup(html_content, 'html.parser')

# Find all h2 elements with the target class
news_headings = soup.find_all('h2', class_='top-sec-title')

for heading in news_headings:
    # Get the parent <a> tag of the h2
    parent_link = heading.parent
    # Extract the href attribute from the a tag
    article_url = parent_link.get('href')
    # Extract the cleaned text from the h2
    article_title = heading.get_text(strip=True)
    
    print(f"Article URL: {article_url}")
    print(f"Article Title: {article_title}")
    print("---")

Key Notes:

  • We use class_='top-sec-title' instead of attrs={'class': ...} here (both work, but class_ is a cleaner shorthand in BeautifulSoup for targeting class attributes, since class is a reserved keyword in Python).
  • heading.parent directly gives us the immediate parent element of the <h2>, which is the <a> tag we need.
  • get_text(strip=True) removes any extra whitespace (like line breaks or leading/trailing spaces) from the heading text, making it cleaner than just .text.

This approach is more reliable for your new HTML structure because the <a> tag doesn't have a unique class to target directly—instead, the <h2> does, so starting there makes sense.

内容的提问来源于stack exchange,提问作者Tayyab Nasir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:59:42