Python BS4爬取网站游戏标题时连带提取span文本该如何解决
解决方法
首先分析无效代码的问题:
bs4.find_all()返回的是h3标签的集合(列表),无法直接对列表调用.find_all()或.text属性,这是之前尝试的代码失效的核心原因。- 每个h3标签内部都包含了带序号的div子标签,直接提取h3的文本会同时拿到序号和标题内容。
- 原代码使用
all作为变量名,和Python内置函数重名,可能引发潜在问题。
修复后的完整代码
import requests from bs4 import BeautifulSoup r = requests.get("https://www.pocketgamer.com/android/best-horror-games/?page=1", headers= {'User-agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:61.0) Gecko/20100101 Firefox/61.0'}) c = r.content bs4 = BeautifulSoup(c,"html.parser") all_h3 = bs4.find_all("h3",{"class":"indent"}) for h3 in all_h3: # 移除h3内部带序号的div标签 if h3.div: h3.div.decompose() # 提取纯文本并去掉首尾空白 game_title = h3.text.strip() if game_title: print(game_title)
运行效果
代码运行后会直接输出纯净的游戏标题,结果如下:
Fran Bow Bendy and the Ink Machine Five Nights at Freddy's Sanitarium OXENFREE Thimbleweed Park Samsara Room
内容的提问来源于stack exchange,提问作者CoderToBeWon
相关产品推荐
相关产品推荐

