You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Beautiful Soup提取每个Div中的首个Href链接?

Extracting Href URLs from MetroLyrics Top 100

Your current code is iterating through unnecessary nested elements and not specifically targeting the anchor (<a>) tags that hold the URLs you need. Let's simplify this to get straight to the href values you're after.

Corrected Code

import bs4 as bs
import urllib.request

masterURL = 'http://www.metrolyrics.com/top100.html'
sauce = urllib.request.urlopen(masterURL).read()
soup = bs.BeautifulSoup(sauce, 'lxml')

# Target all anchor tags inside the song-list unordered lists
for link in soup.select('ul.song-list li a'):
    # Extract the 'href' attribute from each anchor tag
    song_url = link.get('href')
    if song_url:  # Optional: Skip any links without a valid href
        print(song_url)

Key Changes Explained:

  • Using soup.select(): This method uses CSS selectors to directly target the elements you care about. The selector ul.song-list li a finds all <a> tags nested inside <li> elements within <ul> tags with the class song-list—exactly the structure holding the song URLs.
  • Accessing the href attribute: Instead of printing full element content, link.get('href') pulls just the URL value from the anchor tag's href attribute.
  • Optional validity check: The if song_url line ensures we don't print empty or missing href values (a safeguard in case some elements lack links, though unlikely here).

This approach is cleaner, more efficient, and directly returns the highlighted URLs you want without extra nested loops.

内容的提问来源于stack exchange,提问作者Ian-Fogelman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:35:41