Python歌词爬虫执行后反复循环处理首个歌手,如何排查解决?
问题原因分析
- 缩进逻辑错误:遍历专辑列表时,只有符合href匹配规则的专辑才会重新赋值
albumsoup、albumname变量,但后续创建专辑目录、拉取歌曲列表、处理歌曲的代码没有缩进进if判断块内。如果遍历到不符合规则的专辑,代码会复用上一个符合条件的专辑的albumsoup和albumname,重复处理该专辑的所有歌曲,导致无限循环卡在第一个歌手的处理流程中。 - 历史记录不同步:程序启动时会读取本地hist文件到内存的
history列表,但后续处理完新歌写入hist文件后,没有同步把新的歌曲href追加到内存的history列表里。当重复遍历到同一首歌时,判断if song['href'] in history会返回False,导致同一首歌被重复处理,程序无法推进。 - 缺少异常捕获:所有的网络请求、页面元素解析逻辑都没有加异常捕获,一旦触发反爬、请求超时、页面结构变动导致元素找不到等问题,程序会直接抛出异常终止,无法继续处理后续歌手。
- 冗余代码:开头请求的站点首页
frontpage以及解析得到的soupfront全程没有被使用,属于无用请求。
修正后代码
import os from bs4 import BeautifulSoup import ssl import time os.chdir("D:/Folder") import urllib.request if os.path.isfile('hist'): with open('hist', 'r', encoding='utf-8') as file: history = file.read().split() else: history=[] artists=["lil wayne","bob dylan","beyonce"] ssl._create_default_https_context = ssl._create_unverified_context urlhome = "https://www.lyricsfreak.com/" for artist in artists: artist_path = "D:/Folder/"+str(artist) if not os.path.exists(artist_path): os.mkdir(artist_path) link=urlhome+str(artist[0])+"/"+artist.replace(" ","+") try: getartist=urllib.request.urlopen(link, timeout=10) except Exception as e: print(f"请求歌手{artist}页面失败:{e},跳过该歌手") continue artistpage = BeautifulSoup(getartist,features="lxml") albums=artistpage.findAll("a", attrs={"class":"lf-link lf-link--secondary"}) for album in albums: if str(artist[0])+"/"+artist.replace(" ","+") in album["href"]: albumurl = "https://www.lyricsfreak.com"+album["href"] try: albumpage = urllib.request.urlopen(albumurl, timeout=10) except Exception as e: print(f"请求专辑{album.text}页面失败:{e},跳过该专辑") continue albumsoup = BeautifulSoup(albumpage,features="lxml") # 兼容年份不存在的情况 try: albumyear = albumsoup.find("div",attrs={"class":"lf-album__meta-item"}).text.strip()[-6:] except: albumyear = "未知年份" albumname = album.text.strip()+" "+albumyear album_path = artist_path + "/" + albumname if not os.path.exists(album_path): os.mkdir(album_path) songs = albumsoup.findAll("a",href=True,attrs={"class":"lf-link lf-link--secondary"}) for song in songs: if song['href'] in history: print('Skipping', song['href'], '-already on drive') continue time.sleep(3) if "/album/" not in song["href"]: songurl = "https://www.lyricsfreak.com"+song["href"] try: songpage = urllib.request.urlopen(songurl, timeout=10) except Exception as e: print(f"请求歌曲{song.text}页面失败:{e},跳过该歌曲") continue songsoup = BeautifulSoup(songpage,features="lxml") try: songname = songsoup.find("span",attrs={"class":"item-header-color"}).text[:-7] lyrics = songsoup.find("div",attrs={"id":"content"}) fixedlyrics = lyrics.text.strip() except Exception as e: print(f"解析歌曲{song.text}内容失败:{e},跳过该歌曲") continue # 写入歌词时指定编码避免乱码 with open(album_path + "/" + (songname)+".txt","w", encoding="utf-8") as lyricfile: lyricfile.write(fixedlyrics) with open('hist', 'a', encoding='utf-8') as file: file.write(song['href'] + '\n') # 同步更新内存中的历史记录 history.append(song['href']) print("parsing "+str(songname))
内容的提问来源于stack exchange,提问作者Youngun
相关产品推荐
相关产品推荐

