如何使用BeautifulSoup提取Artist results与Song results的目标内容?
Solution to Extract Both Artist and Song Results
Got it, let's fix this so you can grab both artist and song results at once! Based on your description of the nested <div class="panel"> structure, here's a straightforward approach:
Step-by-Step Explanation
Your current code only targets the "Artist results" panel, but we can iterate through all panels on the page, check their heading to categorize the content, then extract the links accordingly.
Working Code Example
import re from bs4 import BeautifulSoup # Assume your soup object is already created from the page content artist_results = [] song_results = [] # Get all panel containers on the page all_panels = soup.find_all('div', class_='panel') for panel in all_panels: # Grab the panel heading to identify what type of results it holds panel_heading = panel.find('div', class_='panel-heading') if not panel_heading: continue # Skip panels without a heading heading_text = panel_heading.get_text(strip=True) # Check if it's Artist results if re.search(r'Artist results:', heading_text): # Extract all artist names from the links artists = [a.get_text(strip=True) for a in panel.find_all('a')] artist_results.extend(artists) # Check if it's Song results elif re.search(r'Song results:', heading_text): # Extract full song entries (like "ASAP by T.I") songs = [a.get_text(strip=True) for a in panel.find_all('a')] song_results.extend(songs) # Print out the collected results print("=== Artist Results ===") for artist in artist_results: print(artist) print("\n=== Song Results ===") for song in song_results: print(song)
Key Improvements
- We use
soup.find_all('div', class_='panel')to get all result panels instead of just one. - For each panel, we check its heading to determine if it's artist or song content.
- We store the results in separate lists so you can work with them individually later.
Notes
- If the heading text varies slightly (e.g., extra spaces), the regex check (
re.search) will still match it reliably. - If you need the actual URLs instead of just the text, replace
a.get_text(strip=True)witha['href'].
内容的提问来源于stack exchange,提问作者hsel
相关产品推荐
相关产品推荐

