从SoundCloud搜索结果HTML提取指定用户ID的技术求助
Let's break down why your code isn't returning results and fix it step by step:
1. Fix the Invalid Request URL
Your URL has an accidental space before filter.created_at, which makes the request target a broken page. Remove that extra space to hit the correct search endpoint:
import requests import re from bs4 import BeautifulSoup url = "https://soundcloud.com/search/sounds?q=f&filter.created_at=last_year&filter.genre_or_tag=hip-hop%20%26%20rap" html = requests.get(url)
2. Correct the Class Name Typo
You're targeting sound_coverArt (single underscore), but the actual HTML class is sound__coverArt (double underscore). This mismatch means BeautifulSoup can't find any matching elements. Update your selector:
soup = BeautifulSoup(html.text, 'html.parser') for sound_link in soup.findAll("a", {"class": "sound__coverArt"}): print(sound_link.get('href'))
3. Address Dynamic Content Loading (Critical Fix)
Even with the above fixes, you might still get no results. SoundCloud loads search results dynamically using JavaScript—requests only fetches the initial static HTML, which doesn't include the actual sound listings.
To pull in the dynamically loaded content, use selenium to simulate a real browser that renders the page fully:
First, install the required package:
pip install selenium
Then use this updated code (make sure you have ChromeDriver installed and matched to your Chrome browser version):
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # Initialize browser driver driver = webdriver.Chrome() url = "https://soundcloud.com/search/sounds?q=f&filter.created_at=last_year&filter.genre_or_tag=hip-hop%20%26%20rap" driver.get(url) # Wait for sound elements to load, then scroll to fetch more results WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "sound__coverArt")) ) # Scroll to bottom to load additional results (repeat this block if you need more) driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) # Give time for new content to render # Extract and isolate user IDs from the href attributes sound_links = driver.find_elements(By.CLASS_NAME, "sound__coverArt") for link in sound_links: href = link.get_attribute('href') if href: # Split the href to get the user ID (format: /user-id/track-name) user_id = href.split('/')[1] print(user_id) driver.quit()
4. Bonus: Isolate User IDs Directly
The code above extracts just the user ID from the full href string, which is exactly what you need instead of printing the entire link.
Quick Notes:
- Match your ChromeDriver version to your installed Chrome browser to avoid compatibility issues.
- Adjust the scroll/sleep timing if you need to load more results.
- Be mindful of SoundCloud's rate limits and anti-scraping measures—don't overload their servers with rapid requests.
内容的提问来源于stack exchange,提问作者BohdanS

