You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python歌词爬虫执行后反复循环处理首个歌手,如何排查解决?

问题原因分析
  • 缩进逻辑错误:遍历专辑列表时,只有符合href匹配规则的专辑才会重新赋值albumsoup、albumname变量,但后续创建专辑目录、拉取歌曲列表、处理歌曲的代码没有缩进进if判断块内。如果遍历到不符合规则的专辑,代码会复用上一个符合条件的专辑的albumsoup和albumname,重复处理该专辑的所有歌曲,导致无限循环卡在第一个歌手的处理流程中。
  • 历史记录不同步:程序启动时会读取本地hist文件到内存的history列表,但后续处理完新歌写入hist文件后,没有同步把新的歌曲href追加到内存的history列表里。当重复遍历到同一首歌时,判断if song['href'] in history会返回False,导致同一首歌被重复处理,程序无法推进。
  • 缺少异常捕获:所有的网络请求、页面元素解析逻辑都没有加异常捕获,一旦触发反爬、请求超时、页面结构变动导致元素找不到等问题,程序会直接抛出异常终止,无法继续处理后续歌手。
  • 冗余代码:开头请求的站点首页frontpage以及解析得到的soupfront全程没有被使用,属于无用请求。
修正后代码
import os
from bs4 import BeautifulSoup
import ssl
import time
os.chdir("D:/Folder")

import urllib.request

if os.path.isfile('hist'):    
    with open('hist', 'r', encoding='utf-8') as file:
        history = file.read().split()
else:
    history=[]
    
artists=["lil wayne","bob dylan","beyonce"]
ssl._create_default_https_context = ssl._create_unverified_context
urlhome = "https://www.lyricsfreak.com/"

for artist in artists:
    artist_path = "D:/Folder/"+str(artist)
    if not os.path.exists(artist_path):
        os.mkdir(artist_path)
    link=urlhome+str(artist[0])+"/"+artist.replace(" ","+")
    try:
        getartist=urllib.request.urlopen(link, timeout=10)
    except Exception as e:
        print(f"请求歌手{artist}页面失败:{e},跳过该歌手")
        continue
    artistpage = BeautifulSoup(getartist,features="lxml")
    albums=artistpage.findAll("a", attrs={"class":"lf-link lf-link--secondary"})

    for album in albums:
        if str(artist[0])+"/"+artist.replace(" ","+") in album["href"]:
            albumurl = "https://www.lyricsfreak.com"+album["href"]
            try:
                albumpage = urllib.request.urlopen(albumurl, timeout=10)
            except Exception as e:
                print(f"请求专辑{album.text}页面失败:{e},跳过该专辑")
                continue
            albumsoup = BeautifulSoup(albumpage,features="lxml")
            # 兼容年份不存在的情况
            try:
                albumyear = albumsoup.find("div",attrs={"class":"lf-album__meta-item"}).text.strip()[-6:]
            except:
                albumyear = "未知年份"
            albumname = album.text.strip()+" "+albumyear
            album_path = artist_path + "/" + albumname
            if not os.path.exists(album_path):
                os.mkdir(album_path)
            songs = albumsoup.findAll("a",href=True,attrs={"class":"lf-link lf-link--secondary"})
            
            for song in songs:
                if song['href'] in history:
                    print('Skipping', song['href'], '-already on drive')
                    continue

                time.sleep(3)
                if "/album/" not in song["href"]:
                    songurl = "https://www.lyricsfreak.com"+song["href"]
                    try:
                        songpage = urllib.request.urlopen(songurl, timeout=10)
                    except Exception as e:
                        print(f"请求歌曲{song.text}页面失败:{e},跳过该歌曲")
                        continue
                    songsoup = BeautifulSoup(songpage,features="lxml")
                    try:
                        songname = songsoup.find("span",attrs={"class":"item-header-color"}).text[:-7]
                        lyrics = songsoup.find("div",attrs={"id":"content"})
                        fixedlyrics = lyrics.text.strip()
                    except Exception as e:
                        print(f"解析歌曲{song.text}内容失败:{e},跳过该歌曲")
                        continue
                    # 写入歌词时指定编码避免乱码
                    with open(album_path + "/" + (songname)+".txt","w", encoding="utf-8") as lyricfile:
                        lyricfile.write(fixedlyrics)
                    with open('hist', 'a', encoding='utf-8') as file:
                        file.write(song['href'] + '\n')
                    # 同步更新内存中的历史记录
                    history.append(song['href'])
                    print("parsing "+str(songname))

内容的提问来源于stack exchange,提问作者Youngun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 17:24:03