如何批量解析搜索后URL的特定字段?(JS/Python优先)
Hey there! Let's walk through how to build this step by step—this is totally approachable even if you're just getting started with web scraping and data processing. Here's your beginner-friendly guide:
一、First, map out the core workflow
Let's break this into simple, actionable steps so you don't get overwhelmed:
- Pull all artist names from your CSV file
- For each name, run a search on Songkick and grab the artist's detail page URL
- Extract the
artist_idfrom that URL (like grabbing301329fromhttps://www.songkick.com/artists/301329-rac) - Save the artist name + matching ID back to a new file (CSV, JSON—whatever works for you)
二、Prep your tools
You'll need a few basic libraries to handle the heavy lifting:
For Python (great for beginners with data tasks):
pandas: Reads/writes CSV files in 2 lines of coderequests: Sends web requests to fetch Songkick pagesBeautifulSoup: Parses the HTML to find the links you need
Install them with:
pip install pandas requests beautifulsoup4
For JavaScript (Node.js environment):
csv-parser/papaparse: Reads CSV dataaxios: Sends web requestscheerio: Parses HTML (like a lightweight jQuery for Node)
Install them after initializing a Node project:
npm init -y npm install axios cheerio csv-parser csv-writer
Quick critical note: Always check Songkick's robots.txt and terms of service before scraping. Don't spam requests—add a 1-2 second delay between calls to avoid getting your IP blocked.
三、Python implementation (step-by-step)
1. Read your CSV
Assuming your CSV has a column named artist_name:
import pandas as pd # Load the CSV and pull out the artist names df = pd.read_csv("your_artists.csv") artist_names = df["artist_name"].tolist()
2. Fetch the artist's detail page URL
We'll hit Songkick's search page, then scrape the first matching artist link:
import requests from bs4 import BeautifulSoup import time def get_artist_detail_url(artist_name): # Encode the artist name for the URL (handles spaces/special chars) encoded_name = requests.utils.quote(artist_name) search_url = f"https://www.songkick.com/search?query={encoded_name}" # Mimic a browser with a User-Agent header to avoid being blocked headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(search_url, headers=headers) if response.status_code != 200: print(f"Failed to search for {artist_name}") return None # Parse the HTML to find the first artist link soup = BeautifulSoup(response.text, "html.parser") artist_link = soup.find("a", class_="artist-link") if artist_link: return artist_link["href"] # Returns something like "/artists/301329-rac" else: print(f"No artist found for {artist_name}") return None
3. Extract the artist_id
Use a regex to pull the numeric ID from the URL path:
import re def extract_artist_id(url_path): # Match the number between "/artists/" and the hyphen match = re.search(r"/artists/(\d+)-", url_path) return match.group(1) if match else None
4. Batch process and save results
Tie it all together, add delays, and write to a new CSV:
results = [] for name in artist_names: url_path = get_artist_detail_url(name) if url_path: artist_id = extract_artist_id(url_path) results.append({"artist_name": name, "artist_id": artist_id}) # Wait 1 second between requests to be polite time.sleep(1) # Save the final data to a new CSV result_df = pd.DataFrame(results) result_df.to_csv("artists_with_ids.csv", index=False)
四、JavaScript (Node.js) implementation (step-by-step)
1. Read your CSV
const fs = require("fs"); const csv = require("csv-parser"); async function loadArtistNames(csvPath) { const artists = []; return new Promise((resolve, reject) => { fs.createReadStream(csvPath) .pipe(csv()) .on("data", (row) => artists.push(row.artist_name)) .on("end", () => resolve(artists)) .on("error", reject); }); }
2. Fetch the artist's detail page URL
const axios = require("axios"); const cheerio = require("cheerio"); async function getArtistDetailUrl(artistName) { const encodedName = encodeURIComponent(artistName); const searchUrl = `https://www.songkick.com/search?query=${encodedName}`; const headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" }; try { const response = await axios.get(searchUrl, { headers }); const $ = cheerio.load(response.data); // Grab the first artist link return $(".artist-link").first().attr("href"); } catch (err) { console.log(`Failed to search for ${artistName}: ${err.message}`); return null; } }
3. Extract the artist_id
function extractArtistId(urlPath) { const match = urlPath.match(/\/artists\/(\d+)-/); return match ? match[1] : null; }
4. Batch process and save results
const createCsvWriter = require("csv-writer").createObjectCsvWriter; async function processAllArtists(inputCsv, outputCsv) { const artists = await loadArtistNames(inputCsv); const results = []; for (const name of artists) { const urlPath = await getArtistDetailUrl(name); if (urlPath) { const artistId = extractArtistId(urlPath); results.push({ artist_name: name, artist_id: artistId }); } // Add a 1-second delay await new Promise(resolve => setTimeout(resolve, 1000)); } // Write results to a new CSV const csvWriter = createCsvWriter({ path: outputCsv, header: [ { id: "artist_name", title: "artist_name" }, { id: "artist_id", title: "artist_id" } ] }); await csvWriter.writeRecords(results); console.log("Done! Results saved to", outputCsv); } // Run the function processAllArtists("your_artists.csv", "artists_with_ids.csv");
五、Beginner pro tips
- Fixing mismatches: Sometimes the first search result isn't the right artist. You can add a check to compare the displayed name on the link with your input name to improve accuracy.
- Retrying failed requests: If some searches fail, add a retry loop (max 2-3 tries) instead of skipping them entirely.
- Avoiding blocks: If you get blocked, switch your
User-Agentstring or try using a free proxy service (start simple first though!).
内容的提问来源于stack exchange,提问作者BS100

