网页抓取:多H3标签下如何定位指定H3内的目标文本?
网页抓取:用BeautifulSoup精准定位特定H3标签文本
问题场景
抓取目标页面时,每个section.yf-work节点下存在两个H3标签:一个是排名(如显示"1"),另一个是嵌套在div.yf-anchor.anchor容器内的书籍标题。现有代码只能获取第一个H3的内容,需要精准提取第二个H3的书籍标题,并整合到主抓取流程中。
解决思路
核心是通过父容器过滤目标H3:目标书籍标题的H3被包裹在class="yf-anchor anchor"的div中,针对每个book节点,先定位该div,再从中提取H3文本即可。
之前代码的错误点:
- CSS选择器写法错误:多个类名需用
.连接,div.yf-anchor anchor h3是无效写法,正确应为div.yf-anchor.anchor h3 select()返回元素列表,不能直接调用.text,用select_one()可直接返回单个目标元素
修正后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd url = "https://www.hebban.nl/rank" # 模拟浏览器请求头 response = requests.get(url, headers={'user-agent':'Mozilla/5.0'}) soup = BeautifulSoup(response.content, 'html.parser') data = [] books = soup.find_all('section', class_='yf-work') for book in books: # 获取排名:优先从指定i标签提取,否则取第一个H3文本 rank_elem = book.find('i', class_='yf-checked fa fa-check-square-o') rank = rank_elem.text.strip() if rank_elem else book.find_all('h3')[0].text.strip() # 精准获取书籍标题:通过父div定位目标H3 anchor_div = book.find('div', class_='yf-anchor anchor') title = anchor_div.find('h3').text.strip() if anchor_div else None # 获取作者信息 author = book.h4.text.strip() if book.find('h4') else None # 获取书籍分类 genre_elem = book.find('a', class_='btn btn4 yf-genre') genre = genre_elem.text.strip() if genre_elem else None # 整理数据存入列表 data.append({ 'rank': rank, 'author': author, 'title': title, 'genres': genre, 'scraped_date': pd.Timestamp.today().strftime('%Y-%m-%d') }) # 生成DataFrame并输出 df = pd.DataFrame(data) print(df)
关键代码说明
book.find('div', class_='yf-anchor anchor'):精准定位包裹书籍标题的父容器,避免误触排名H3- 也可替换为
book.select_one('div.yf-anchor.anchor h3')直接获取目标H3元素,两种方式效果一致 - 所有元素提取都增加了空值判断,避免因页面结构变化导致代码报错
内容的提问来源于stack exchange,提问作者jsb92
相关产品推荐
相关产品推荐

