如何用Python+BeautifulSoup批量提取同class的div中所有URL?
问题:批量提取指定class的div内容并存入Google Sheets
我需要提取网站中所有class为list-pages-item的div内的文本(类名称)和href链接,并存入Google Sheets。但目前代码只能提取第一个div的信息,求用循环或更高效的方式实现批量提取。
网页HTML示例
<div class="list-pages-item"> <p><a href="/abyssal-angel-ac">Abyssal Angel (AC)</a></p> <div class="list-pages-item"> <p><a href="/acolyte">Acolyte</a></p> <div class="list-pages-item"> <p><a href="/alpha-omega-non-legend">Alpha Omega (Non-Legend)</a></p>
现有代码
def parse(soup): #Grabs Classes htmlcode = soup.find('div', class_='list-pages-item') base_url = "http://aqwwiki.wikidot.com" ClassName = htmlcode.text.strip() for a in htmlcode.find_all('a', href=True): wiki = (base_url + a['href']) Outcome = {'Class': ClassName, 'URL': wiki} return Outcome
解决方案
核心修改是把获取单个元素的find()换成获取所有匹配元素的find_all(),再通过循环逐个处理每个div:
修改后的代码
def parse(soup): # 批量抓取所有类信息 items = soup.find_all('div', class_='list-pages-item') base_url = "http://aqwwiki.wikidot.com" outcomes = [] for item in items: # 提取类名称 class_name = item.text.strip() # 提取链接(每个div仅含一个a标签,直接用find即可) a_tag = item.find('a', href=True) if a_tag: # 容错处理:避免无a标签的div导致报错 wiki_url = base_url + a_tag['href'] outcomes.append({'Class': class_name, 'URL': wiki_url}) return outcomes
关键说明
soup.find_all()会返回页面中所有符合list-pages-item类的div元素列表,而非单个元素- 循环遍历每个div,分别提取文本内容和链接,将每个条目封装为字典后加入结果列表
- 加入
if a_tag:判断,防止页面中存在不含a标签的目标div时触发报错 - 最终返回的是包含所有条目信息的列表,方便后续批量写入Google Sheets(可通过
gspread等库实现写入操作)
内容的提问来源于stack exchange,提问作者JonKodirovaniye
相关产品推荐
相关产品推荐

