You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量解析搜索后URL的特定字段?(JS/Python优先)

Hey there! Let's walk through how to build this step by step—this is totally approachable even if you're just getting started with web scraping and data processing. Here's your beginner-friendly guide:

入门思路与分步实现指导

一、First, map out the core workflow

Let's break this into simple, actionable steps so you don't get overwhelmed:

  • Pull all artist names from your CSV file
  • For each name, run a search on Songkick and grab the artist's detail page URL
  • Extract the artist_id from that URL (like grabbing 301329 from https://www.songkick.com/artists/301329-rac)
  • Save the artist name + matching ID back to a new file (CSV, JSON—whatever works for you)

二、Prep your tools

You'll need a few basic libraries to handle the heavy lifting:

For Python (great for beginners with data tasks):

  • pandas: Reads/writes CSV files in 2 lines of code
  • requests: Sends web requests to fetch Songkick pages
  • BeautifulSoup: Parses the HTML to find the links you need
    Install them with:
pip install pandas requests beautifulsoup4

For JavaScript (Node.js environment):

  • csv-parser/papaparse: Reads CSV data
  • axios: Sends web requests
  • cheerio: Parses HTML (like a lightweight jQuery for Node)
    Install them after initializing a Node project:
npm init -y
npm install axios cheerio csv-parser csv-writer

Quick critical note: Always check Songkick's robots.txt and terms of service before scraping. Don't spam requests—add a 1-2 second delay between calls to avoid getting your IP blocked.

三、Python implementation (step-by-step)

1. Read your CSV

Assuming your CSV has a column named artist_name:

import pandas as pd

# Load the CSV and pull out the artist names
df = pd.read_csv("your_artists.csv")
artist_names = df["artist_name"].tolist()

2. Fetch the artist's detail page URL

We'll hit Songkick's search page, then scrape the first matching artist link:

import requests
from bs4 import BeautifulSoup
import time

def get_artist_detail_url(artist_name):
    # Encode the artist name for the URL (handles spaces/special chars)
    encoded_name = requests.utils.quote(artist_name)
    search_url = f"https://www.songkick.com/search?query={encoded_name}"
    
    # Mimic a browser with a User-Agent header to avoid being blocked
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    response = requests.get(search_url, headers=headers)
    if response.status_code != 200:
        print(f"Failed to search for {artist_name}")
        return None
    
    # Parse the HTML to find the first artist link
    soup = BeautifulSoup(response.text, "html.parser")
    artist_link = soup.find("a", class_="artist-link")
    if artist_link:
        return artist_link["href"]  # Returns something like "/artists/301329-rac"
    else:
        print(f"No artist found for {artist_name}")
        return None

3. Extract the artist_id

Use a regex to pull the numeric ID from the URL path:

import re

def extract_artist_id(url_path):
    # Match the number between "/artists/" and the hyphen
    match = re.search(r"/artists/(\d+)-", url_path)
    return match.group(1) if match else None

4. Batch process and save results

Tie it all together, add delays, and write to a new CSV:

results = []
for name in artist_names:
    url_path = get_artist_detail_url(name)
    if url_path:
        artist_id = extract_artist_id(url_path)
        results.append({"artist_name": name, "artist_id": artist_id})
    # Wait 1 second between requests to be polite
    time.sleep(1)

# Save the final data to a new CSV
result_df = pd.DataFrame(results)
result_df.to_csv("artists_with_ids.csv", index=False)

四、JavaScript (Node.js) implementation (step-by-step)

1. Read your CSV

const fs = require("fs");
const csv = require("csv-parser");

async function loadArtistNames(csvPath) {
    const artists = [];
    return new Promise((resolve, reject) => {
        fs.createReadStream(csvPath)
            .pipe(csv())
            .on("data", (row) => artists.push(row.artist_name))
            .on("end", () => resolve(artists))
            .on("error", reject);
    });
}

2. Fetch the artist's detail page URL

const axios = require("axios");
const cheerio = require("cheerio");

async function getArtistDetailUrl(artistName) {
    const encodedName = encodeURIComponent(artistName);
    const searchUrl = `https://www.songkick.com/search?query=${encodedName}`;
    const headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    };

    try {
        const response = await axios.get(searchUrl, { headers });
        const $ = cheerio.load(response.data);
        // Grab the first artist link
        return $(".artist-link").first().attr("href");
    } catch (err) {
        console.log(`Failed to search for ${artistName}: ${err.message}`);
        return null;
    }
}

3. Extract the artist_id

function extractArtistId(urlPath) {
    const match = urlPath.match(/\/artists\/(\d+)-/);
    return match ? match[1] : null;
}

4. Batch process and save results

const createCsvWriter = require("csv-writer").createObjectCsvWriter;

async function processAllArtists(inputCsv, outputCsv) {
    const artists = await loadArtistNames(inputCsv);
    const results = [];

    for (const name of artists) {
        const urlPath = await getArtistDetailUrl(name);
        if (urlPath) {
            const artistId = extractArtistId(urlPath);
            results.push({ artist_name: name, artist_id: artistId });
        }
        // Add a 1-second delay
        await new Promise(resolve => setTimeout(resolve, 1000));
    }

    // Write results to a new CSV
    const csvWriter = createCsvWriter({
        path: outputCsv,
        header: [
            { id: "artist_name", title: "artist_name" },
            { id: "artist_id", title: "artist_id" }
        ]
    });

    await csvWriter.writeRecords(results);
    console.log("Done! Results saved to", outputCsv);
}

// Run the function
processAllArtists("your_artists.csv", "artists_with_ids.csv");

五、Beginner pro tips

  • Fixing mismatches: Sometimes the first search result isn't the right artist. You can add a check to compare the displayed name on the link with your input name to improve accuracy.
  • Retrying failed requests: If some searches fail, add a retry loop (max 2-3 tries) instead of skipping them entirely.
  • Avoiding blocks: If you get blocked, switch your User-Agent string or try using a free proxy service (start simple first though!).

内容的提问来源于stack exchange,提问作者BS100

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:27:59