You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从SoundCloud搜索结果HTML提取指定用户ID的技术求助

Fixing Your SoundCloud User ID Scraper

Let's break down why your code isn't returning results and fix it step by step:

1. Fix the Invalid Request URL

Your URL has an accidental space before filter.created_at, which makes the request target a broken page. Remove that extra space to hit the correct search endpoint:

import requests
import re
from bs4 import BeautifulSoup

url = "https://soundcloud.com/search/sounds?q=f&filter.created_at=last_year&filter.genre_or_tag=hip-hop%20%26%20rap"
html = requests.get(url)

2. Correct the Class Name Typo

You're targeting sound_coverArt (single underscore), but the actual HTML class is sound__coverArt (double underscore). This mismatch means BeautifulSoup can't find any matching elements. Update your selector:

soup = BeautifulSoup(html.text, 'html.parser')
for sound_link in soup.findAll("a", {"class": "sound__coverArt"}):
    print(sound_link.get('href'))

3. Address Dynamic Content Loading (Critical Fix)

Even with the above fixes, you might still get no results. SoundCloud loads search results dynamically using JavaScript—requests only fetches the initial static HTML, which doesn't include the actual sound listings.

To pull in the dynamically loaded content, use selenium to simulate a real browser that renders the page fully:

First, install the required package:

pip install selenium

Then use this updated code (make sure you have ChromeDriver installed and matched to your Chrome browser version):

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# Initialize browser driver
driver = webdriver.Chrome()
url = "https://soundcloud.com/search/sounds?q=f&filter.created_at=last_year&filter.genre_or_tag=hip-hop%20%26%20rap"
driver.get(url)

# Wait for sound elements to load, then scroll to fetch more results
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, "sound__coverArt"))
)

# Scroll to bottom to load additional results (repeat this block if you need more)
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
time.sleep(3)  # Give time for new content to render

# Extract and isolate user IDs from the href attributes
sound_links = driver.find_elements(By.CLASS_NAME, "sound__coverArt")
for link in sound_links:
    href = link.get_attribute('href')
    if href:
        # Split the href to get the user ID (format: /user-id/track-name)
        user_id = href.split('/')[1]
        print(user_id)

driver.quit()

4. Bonus: Isolate User IDs Directly

The code above extracts just the user ID from the full href string, which is exactly what you need instead of printing the entire link.

Quick Notes:

  • Match your ChromeDriver version to your installed Chrome browser to avoid compatibility issues.
  • Adjust the scroll/sleep timing if you need to load more results.
  • Be mindful of SoundCloud's rate limits and anti-scraping measures—don't overload their servers with rapid requests.

内容的提问来源于stack exchange,提问作者BohdanS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:27:16