如何用Python Beautiful Soup提取每个Div中的首个Href链接?
Extracting Href URLs from MetroLyrics Top 100
Your current code is iterating through unnecessary nested elements and not specifically targeting the anchor (<a>) tags that hold the URLs you need. Let's simplify this to get straight to the href values you're after.
Corrected Code
import bs4 as bs import urllib.request masterURL = 'http://www.metrolyrics.com/top100.html' sauce = urllib.request.urlopen(masterURL).read() soup = bs.BeautifulSoup(sauce, 'lxml') # Target all anchor tags inside the song-list unordered lists for link in soup.select('ul.song-list li a'): # Extract the 'href' attribute from each anchor tag song_url = link.get('href') if song_url: # Optional: Skip any links without a valid href print(song_url)
Key Changes Explained:
- Using
soup.select(): This method uses CSS selectors to directly target the elements you care about. The selectorul.song-list li afinds all<a>tags nested inside<li>elements within<ul>tags with the classsong-list—exactly the structure holding the song URLs. - Accessing the
hrefattribute: Instead of printing full element content,link.get('href')pulls just the URL value from the anchor tag'shrefattribute. - Optional validity check: The
if song_urlline ensures we don't print empty or missing href values (a safeguard in case some elements lack links, though unlikely here).
This approach is cleaner, more efficient, and directly returns the highlighted URLs you want without extra nested loops.
内容的提问来源于stack exchange,提问作者Ian-Fogelman
相关产品推荐
相关产品推荐

