You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取Artist results与Song results的目标内容?

Solution to Extract Both Artist and Song Results

Got it, let's fix this so you can grab both artist and song results at once! Based on your description of the nested <div class="panel"> structure, here's a straightforward approach:

Step-by-Step Explanation

Your current code only targets the "Artist results" panel, but we can iterate through all panels on the page, check their heading to categorize the content, then extract the links accordingly.

Working Code Example

import re
from bs4 import BeautifulSoup

# Assume your soup object is already created from the page content
artist_results = []
song_results = []

# Get all panel containers on the page
all_panels = soup.find_all('div', class_='panel')

for panel in all_panels:
    # Grab the panel heading to identify what type of results it holds
    panel_heading = panel.find('div', class_='panel-heading')
    if not panel_heading:
        continue  # Skip panels without a heading
    
    heading_text = panel_heading.get_text(strip=True)
    
    # Check if it's Artist results
    if re.search(r'Artist results:', heading_text):
        # Extract all artist names from the links
        artists = [a.get_text(strip=True) for a in panel.find_all('a')]
        artist_results.extend(artists)
    # Check if it's Song results
    elif re.search(r'Song results:', heading_text):
        # Extract full song entries (like "ASAP by T.I")
        songs = [a.get_text(strip=True) for a in panel.find_all('a')]
        song_results.extend(songs)

# Print out the collected results
print("=== Artist Results ===")
for artist in artist_results:
    print(artist)

print("\n=== Song Results ===")
for song in song_results:
    print(song)

Key Improvements

  • We use soup.find_all('div', class_='panel') to get all result panels instead of just one.
  • For each panel, we check its heading to determine if it's artist or song content.
  • We store the results in separate lists so you can work with them individually later.

Notes

  • If the heading text varies slightly (e.g., extra spaces), the regex check (re.search) will still match it reliably.
  • If you need the actual URLs instead of just the text, replace a.get_text(strip=True) with a['href'].

内容的提问来源于stack exchange,提问作者hsel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:42:15