Java新手求助:使用JSoup爬取choon.co艺术家名称无结果
Hey there! No need to apologize—we all start somewhere, and let's get your scraper working step by step. The main issues here are a small syntax mistake in your Jsoup selector, plus a need to tweak how we target the artist elements for precision.
1. Fix the Class Selector Syntax
Your current selector doc.select("class.ch-track-list__cell") is incorrect. In Jsoup (and standard CSS), to select elements by their class name, you use a dot (.) prefix instead of writing class.. The correct selector should be ".ch-track-list__cell".
2. Target the Artist Link Precisely
Looking at the target page's structure, the artist name is nested inside a specific child element of the track list cell. Grabbing the first <a> tag might not always target the artist (it could pick up other links in the cell). Instead, use a more specific selector to directly target the artist-specific cell and its link: .ch-track-list__cell.ch-track-list__cell--artist a.
Corrected Code
Here's your updated code with these fixes, plus extra debugging clarity:
package com.manu.scraper; import java.io.IOException; import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.jsoup.select.Elements; public class Spider { public static void main(String[] args) { System.out.println("Fetching..."); try { // Use a modern user agent to avoid being blocked by the site Document doc = Jsoup.connect("https://www.choon.co/playlists/genre_12-bar-blues/12-bar-blues") .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .get(); // Select only the anchor tags inside artist-specific track list cells Elements artistLinks = doc.select(".ch-track-list__cell.ch-track-list__cell--artist a"); int i = 0; for(Element artist : artistLinks) { i++; System.out.println(i + " " + artist.text()); } // Add a fallback message if no artists are found if (artistLinks.isEmpty()) { System.out.println("No artists found—double-check the selector or page structure!"); } System.out.println("Done!"); } catch (IOException e) { e.printStackTrace(); } } }
Additional Notes
- User Agent Update: I swapped your user agent for a more recent Chrome version to reduce the chance of the site blocking your request.
- Empty Check: Added a check to notify you if no elements are found, which helps debug if the page structure changes later.
- Dynamic Content Edge Case: If this still doesn't work, the site might load tracks dynamically with JavaScript. In that case, you'd need a tool like Selenium instead of Jsoup—but let's start with the fixes above first, as they should resolve the immediate issue.
内容的提问来源于stack exchange,提问作者Schaedel420

