如何用Python+BeautifulSoup获取网站全页面列表?附现有URL片段
Alright, let's walk through how to scrape all of Ed Sheeran's lyric pages from LyricsFreak using Python and BeautifulSoup—building on the partial URL list you already have. Here's a practical, step-by-step approach:
Step 1: Collect All Song URLs First
Your existing URLs are for individual tracks, but to get every single one, we need to start at Ed Sheeran's artist page (that's where all his song links are listed). We'll also handle pagination in case his discography spans multiple pages.
First, import the necessary libraries:
import requests from bs4 import BeautifulSoup import time from urllib.parse import urljoin
Then, write a function to fetch all song URLs:
def get_all_ed_sheeran_songs(base_url="http://www.lyricsfreak.com/e/ed+sheeran/"): song_urls = set() # Using a set to automatically avoid duplicate URLs current_url = base_url while True: # Add a 2-second delay to avoid overwhelming the server time.sleep(2) # Send a request with a proper User-Agent to mimic a real browser response = requests.get( current_url, headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"} ) if response.status_code != 200: print(f"Oops, couldn't load {current_url} (status code: {response.status_code})") break soup = BeautifulSoup(response.text, "html.parser") # Extract all song links from the page (adjust the selector if the site updates) song_links = soup.select("div.song-list > ul > li > a") for link in song_links: full_url = urljoin(base_url, link["href"]) song_urls.add(full_url) # Check if there's a "Next Page" button to continue scraping next_page = soup.select_one("a.next-page") if not next_page: break # No more pages, exit the loop current_url = urljoin(base_url, next_page["href"]) # Add your existing URLs to the set to make sure none are missing your_existing_urls = [ 'http://www.lyricsfreak.com/e/ed+sheeran/thinking+out+loud_21083784.html', 'http://www.lyricsfreak.com/e/ed+sheeran/photograph_21058341.html', 'http://www.lyricsfreak.com/e/ed+sheeran/a+team_20983411.html', 'http://www.lyricsfreak.com/e/ed+sheeran/i+see+fire_21071421.html', 'http://www.lyricsfreak.com/e/ed+sheeran/perfect_21113253.html' ] song_urls.update(your_existing_urls) return list(song_urls)
Step 2: Scrape Lyrics from Each Song Page
Now that we have all the URLs, let's write a function to pull the lyrics from each page:
def scrape_song_lyrics(song_url): time.sleep(2) response = requests.get( song_url, headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"} ) if response.status_code != 200: print(f"Failed to load lyrics for {song_url}") return None soup = BeautifulSoup(response.text, "html.parser") # Grab the lyrics container (again, adjust the selector if the site changes) lyrics_div = soup.select_one("div.lyrictxt") if lyrics_div: # Extract clean text from the lyrics div lyrics = lyrics_div.get_text(strip=True, separator="\n") song_title = soup.select_one("h1").get_text(strip=True) return { "title": song_title, "url": song_url, "lyrics": lyrics } else: print(f"Couldn't find lyrics for {song_url}") return None
Step 3: Put It All Together
Call the functions to collect URLs and scrape lyrics, then save the results to a file:
if __name__ == "__main__": # Get all song URLs all_songs = get_all_ed_sheeran_songs() print(f"Found {len(all_songs)} total song URLs") # Scrape lyrics for each song all_lyrics_data = [] for url in all_songs: lyric_info = scrape_song_lyrics(url) if lyric_info: all_lyrics_data.append(lyric_info) # Save the results to a text file with open("ed_sheeran_lyrics.txt", "w", encoding="utf-8") as file: for entry in all_lyrics_data: file.write(f"Title: {entry['title']}\n") file.write(f"URL: {entry['url']}\n") file.write("Lyrics:\n") file.write(f"{entry['lyrics']}\n") file.write("\n" + "="*50 + "\n\n") print("Done! All lyrics saved to ed_sheeran_lyrics.txt")
Key Things to Keep in Mind
- Respect the Website: Always check the site's
robots.txtto confirm scraping is allowed. Adjust thetime.sleep()delay if you get blocked—slower is safer. - Selector Updates: If the code stops working, the site probably changed its HTML structure. Use your browser's dev tools to inspect elements and update the CSS selectors (like
div.song-listordiv.lyrictxt). - Duplicates: Using a
setfor URLs ensures you don't scrape the same page twice, even if it appears on multiple pagination pages.
内容的提问来源于stack exchange,提问作者Todd Burus

