You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+BeautifulSoup获取网站全页面列表?附现有URL片段

Alright, let's walk through how to scrape all of Ed Sheeran's lyric pages from LyricsFreak using Python and BeautifulSoup—building on the partial URL list you already have. Here's a practical, step-by-step approach:

Step 1: Collect All Song URLs First

Your existing URLs are for individual tracks, but to get every single one, we need to start at Ed Sheeran's artist page (that's where all his song links are listed). We'll also handle pagination in case his discography spans multiple pages.

First, import the necessary libraries:

import requests
from bs4 import BeautifulSoup
import time
from urllib.parse import urljoin

Then, write a function to fetch all song URLs:

def get_all_ed_sheeran_songs(base_url="http://www.lyricsfreak.com/e/ed+sheeran/"):
    song_urls = set()  # Using a set to automatically avoid duplicate URLs
    current_url = base_url
    
    while True:
        # Add a 2-second delay to avoid overwhelming the server
        time.sleep(2)
        # Send a request with a proper User-Agent to mimic a real browser
        response = requests.get(
            current_url,
            headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}
        )
        
        if response.status_code != 200:
            print(f"Oops, couldn't load {current_url} (status code: {response.status_code})")
            break
        
        soup = BeautifulSoup(response.text, "html.parser")
        
        # Extract all song links from the page (adjust the selector if the site updates)
        song_links = soup.select("div.song-list > ul > li > a")
        for link in song_links:
            full_url = urljoin(base_url, link["href"])
            song_urls.add(full_url)
        
        # Check if there's a "Next Page" button to continue scraping
        next_page = soup.select_one("a.next-page")
        if not next_page:
            break  # No more pages, exit the loop
        current_url = urljoin(base_url, next_page["href"])
    
    # Add your existing URLs to the set to make sure none are missing
    your_existing_urls = [
        'http://www.lyricsfreak.com/e/ed+sheeran/thinking+out+loud_21083784.html',
        'http://www.lyricsfreak.com/e/ed+sheeran/photograph_21058341.html',
        'http://www.lyricsfreak.com/e/ed+sheeran/a+team_20983411.html',
        'http://www.lyricsfreak.com/e/ed+sheeran/i+see+fire_21071421.html',
        'http://www.lyricsfreak.com/e/ed+sheeran/perfect_21113253.html'
    ]
    song_urls.update(your_existing_urls)
    
    return list(song_urls)

Step 2: Scrape Lyrics from Each Song Page

Now that we have all the URLs, let's write a function to pull the lyrics from each page:

def scrape_song_lyrics(song_url):
    time.sleep(2)
    response = requests.get(
        song_url,
        headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}
    )
    
    if response.status_code != 200:
        print(f"Failed to load lyrics for {song_url}")
        return None
    
    soup = BeautifulSoup(response.text, "html.parser")
    # Grab the lyrics container (again, adjust the selector if the site changes)
    lyrics_div = soup.select_one("div.lyrictxt")
    if lyrics_div:
        # Extract clean text from the lyrics div
        lyrics = lyrics_div.get_text(strip=True, separator="\n")
        song_title = soup.select_one("h1").get_text(strip=True)
        return {
            "title": song_title,
            "url": song_url,
            "lyrics": lyrics
        }
    else:
        print(f"Couldn't find lyrics for {song_url}")
        return None

Step 3: Put It All Together

Call the functions to collect URLs and scrape lyrics, then save the results to a file:

if __name__ == "__main__":
    # Get all song URLs
    all_songs = get_all_ed_sheeran_songs()
    print(f"Found {len(all_songs)} total song URLs")
    
    # Scrape lyrics for each song
    all_lyrics_data = []
    for url in all_songs:
        lyric_info = scrape_song_lyrics(url)
        if lyric_info:
            all_lyrics_data.append(lyric_info)
    
    # Save the results to a text file
    with open("ed_sheeran_lyrics.txt", "w", encoding="utf-8") as file:
        for entry in all_lyrics_data:
            file.write(f"Title: {entry['title']}\n")
            file.write(f"URL: {entry['url']}\n")
            file.write("Lyrics:\n")
            file.write(f"{entry['lyrics']}\n")
            file.write("\n" + "="*50 + "\n\n")
    
    print("Done! All lyrics saved to ed_sheeran_lyrics.txt")

Key Things to Keep in Mind

  • Respect the Website: Always check the site's robots.txt to confirm scraping is allowed. Adjust the time.sleep() delay if you get blocked—slower is safer.
  • Selector Updates: If the code stops working, the site probably changed its HTML structure. Use your browser's dev tools to inspect elements and update the CSS selectors (like div.song-list or div.lyrictxt).
  • Duplicates: Using a set for URLs ensures you don't scrape the same page twice, even if it appears on multiple pagination pages.

内容的提问来源于stack exchange,提问作者Todd Burus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:13:14