You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中调用requests.get()请求href链接时抛出错误的问题

Hey there! Let's figure out why your requests.get(href) is throwing errors. There are a couple of common issues here with Azlyrics and your scraping logic—let's break them down and fix your code step by step.

1. Azlyrics Blocks Default Requests (Anti-Scraping)

Azlyrics actively blocks requests that don't look like they're coming from a real browser. When you use requests.get() without custom headers, your request gets flagged and returns a 403 Forbidden error. To fix this, add a User-Agent header to mimic a standard browser.

Right now, you're looping through every <a> tag on the search page and picking the first one with https and lyrics/ in the href. This might grab irrelevant links (like ads or navigation) instead of actual search result links. We can narrow this down to target only the valid search result panels.

3. Missing Error Handling

Your code doesn't handle cases where requests fail (e.g., no internet, broken links) or no valid lyric links are found. Adding basic error handling will make your script way more robust.


Here's the fixed, improved version of your code:

import string
import requests
from bs4 import BeautifulSoup

# Mimic a browser request to avoid being blocked by Azlyrics
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

song_name = input('enter song name : ')
# Format the search query correctly
query = string.capwords(song_name).replace(' ', '+')
url = f'https://search.azlyrics.com/search.php?q={query}'

try:
    # Fetch the search page with headers
    res = requests.get(url, headers=headers)
    res.raise_for_status()  # Trigger error for HTTP issues like 403/404
    soup = BeautifulSoup(res.text, 'html.parser')

    # Target only valid lyric links from search results (more precise than all <a> tags)
    lyric_links = soup.select('div.panel div.table a[href*="lyrics/"]')
    
    if not lyric_links:
        print("No lyric links found for that song.")
    else:
        # Grab the first valid lyric link
        lyric_url = lyric_links[0]['href']
        print(f"Fetching lyrics from: {lyric_url}")
        
        # Fetch the lyric page with headers
        lyric_res = requests.get(lyric_url, headers=headers)
        lyric_res.raise_for_status()
        
        # Optional: Extract and print the actual lyrics (since that's likely your end goal!)
        lyric_soup = BeautifulSoup(lyric_res.text, 'html.parser')
        # Azlyrics stores lyrics in a div without class/id, right after a specific comment
        lyric_div = lyric_soup.find('div', attrs={'class': None, 'id': None})
        if lyric_div:
            print("\nLyrics:\n")
            print(lyric_div.get_text(strip=True, separator='\n'))
        else:
            print("Couldn't extract lyrics from the page.")

except requests.exceptions.RequestException as e:
    print(f"An error occurred during the request: {e}")

Key Improvements:

  • Browser-like Headers: The User-Agent header tricks Azlyrics into thinking your request is from a real user, avoiding 403 errors.
  • Precise Link Selection: Using soup.select() with a CSS selector targets only the actual search result links, skipping ads/navigation.
  • Error Handling: try-except blocks catch network issues and HTTP errors, so your script doesn't crash unexpectedly.
  • Lyric Extraction: Added optional code to pull and display the lyrics (since that's probably what you want after getting the link!).

内容的提问来源于stack exchange,提问作者naveen prajapati

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:51:55