网页爬虫技术需求:提取每个h3标题及对应ul下的li文本内容
提取H3标题对应UL列表的爬虫实现方法
针对你遇到的问题,用Python的BeautifulSoup库可以轻松解决,核心思路是每个H3标签的紧邻下一个UL标签就是它对应的列表,具体实现代码如下:
from bs4 import BeautifulSoup # 你的HTML代码(可替换为爬虫获取的页面内容) html_content = ''' <div class = "some div class"> <div class = "some other div"> <h3> first title </h3> <ul> <li>Square No.1477: Rodman Street. NW And Fordham Street NW</li> <li>Square No.1586: Davenport Street. NW And 44th Street NW</li> <li>Square No.1738: Garrison Street. NW And 41St Street NW</li> <li>Square No.2997:Ingraham Street. NW And Georgia Avenue NW</li> <li>Square No.3145: Illinois Avenue NW, And Decatur Street NW</li> <li>Square No.3292: Madison Street. NW And 3Rd Place NW</li> <li>Square No.3337: 2nd Street. NW And Oglethorpe Street NW</li> <li>Square No.3337: Peabody Street. NW And 2nd Place NW</li> <li>Square No.5441: Livingston Street NW, Nevada Avenue. NW And Legation Street NW</li> </ul> <h3>Second title</h3> <ul> <li>Square No.177: Fordham Street NW</li> <li>Square No.186: 44th Street NW</li> <li>Square No.138: 41St Street NW</li> <li>Square No.997:Ingraham Georgia Avenue NW</li> <li>Square No.314: Decatur Street NW</li> <li>Square No.3292: Madison Street. NW And 3Rd Place NW</li> <li>Square No.333: Oglethorpe Street NW</li> <li>Square No.3337: Peabody Street. NW And 2nd Place NW</li> <li>Square No.5441: Livingston Street NW, Nevada Avenue. NW And Legation Street NW</li> </ul> </div> </div> ''' # 解析HTML soup = BeautifulSoup(html_content, 'html.parser') # 遍历所有H3标签 for h3 in soup.find_all('h3'): # 提取标题文本,strip()去除前后空格 title = h3.get_text(strip=True) print(f"--- {title} ---") # 找到当前H3后面紧邻的UL标签 ul = h3.find_next('ul') # 遍历UL下的所有LI标签,提取文本 for li in ul.find_all('li'): print(f"- {li.get_text(strip=True)}") print() # 空行分隔不同标题的内容
关键代码解释
soup.find_all('h3'):获取页面中所有的H3标题标签h3.find_next('ul'):定位当前H3标签之后的第一个UL标签,确保是该标题对应的列表get_text(strip=True):提取标签内的文本,并自动去除前后多余的空格和换行符
运行代码后,输出格式就是每个标题搭配其下所有LI的文本内容,完全符合你的需求。
内容的提问来源于stack exchange,提问作者Shubham Kapadnis
相关产品推荐
相关产品推荐

