Python网页爬虫循环仅返回首个迭代结果问题求助
问题解决:爬虫循环重复抓取首个文章内容
你的问题根源很明确:循环内始终从全局的BeautifulSoup对象中查找元素,每次都取第一个article的内容,而不是基于当前循环迭代的article对象去查找它的子元素。
修改后的代码如下:
import requests import bs4 url = 'https://coreyms.com' response = requests.get(url) response.raise_for_status() schafer = bs4.BeautifulSoup(response.text, 'html.parser') # 遍历每一篇文章 for article in schafer.find_all('article'): # 基于当前article对象查找标题 header = article.select_one('article a').getText() print(header) # 基于当前article对象查找段落 paragraph = article.select_one('div > p').getText() print(paragraph) # 基于当前article对象查找iframe,先判断是否存在 iframe = article.select_one('iframe') if iframe: link = iframe.get('src') vidID = link.split('/')[4].split('?')[0] ytLink = f'https://youtube.com/watch?v={vidID}' print(ytLink) print()
关键修改点:
- 将所有
schafer.select()替换为article.select_one():select_one()会返回当前元素下匹配的第一个子元素,正好对应单篇文章里的标题、段落和视频框架 - 添加了iframe存在性判断:避免部分文章没有视频时出现索引错误
- 去掉了不必要的索引[0]:
select_one()直接返回单个元素,无需再取列表第一个项
内容的提问来源于stack exchange,提问作者Slatercj
相关产品推荐
相关产品推荐

