You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬虫技术需求:提取每个h3标题及对应ul下的li文本内容

提取H3标题对应UL列表的爬虫实现方法

针对你遇到的问题,用Python的BeautifulSoup库可以轻松解决,核心思路是每个H3标签的紧邻下一个UL标签就是它对应的列表,具体实现代码如下:

from bs4 import BeautifulSoup

# 你的HTML代码(可替换为爬虫获取的页面内容)
html_content = '''
<div class = "some div class">
   <div class = "some other div">
     <h3> first title </h3>
    <ul>
        <li>Square No.1477:  Rodman Street. NW And Fordham Street NW</li>
        <li>Square No.1586:  Davenport Street. NW And 44th Street NW</li>
        <li>Square No.1738:  Garrison Street. NW And 41St Street NW</li>
        <li>Square No.2997:Ingraham Street. NW And Georgia Avenue NW</li>
        <li>Square No.3145: Illinois Avenue NW, And Decatur Street NW</li>
        <li>Square No.3292:  Madison Street. NW And 3Rd Place NW</li>
        <li>Square No.3337:  2nd Street. NW And Oglethorpe Street NW</li>
        <li>Square No.3337: Peabody Street. NW And 2nd Place NW</li>
        <li>Square No.5441: Livingston Street NW, Nevada Avenue. NW And Legation Street NW</li>
    </ul>
   <h3>Second title</h3>
    <ul>
        <li>Square No.177:  Fordham Street NW</li>
        <li>Square No.186:  44th Street NW</li>
        <li>Square No.138:  41St Street NW</li>
        <li>Square No.997:Ingraham Georgia Avenue NW</li>
        <li>Square No.314: Decatur Street NW</li>
        <li>Square No.3292:  Madison Street. NW And 3Rd Place NW</li>
        <li>Square No.333:  Oglethorpe Street NW</li>
        <li>Square No.3337: Peabody Street. NW And 2nd Place NW</li>
        <li>Square No.5441: Livingston Street NW, Nevada Avenue. NW And Legation Street NW</li>
    </ul>
   </div>
</div>
'''

# 解析HTML
soup = BeautifulSoup(html_content, 'html.parser')

# 遍历所有H3标签
for h3 in soup.find_all('h3'):
    # 提取标题文本,strip()去除前后空格
    title = h3.get_text(strip=True)
    print(f"--- {title} ---")
    
    # 找到当前H3后面紧邻的UL标签
    ul = h3.find_next('ul')
    # 遍历UL下的所有LI标签,提取文本
    for li in ul.find_all('li'):
        print(f"- {li.get_text(strip=True)}")
    print()  # 空行分隔不同标题的内容

关键代码解释

  • soup.find_all('h3'):获取页面中所有的H3标题标签
  • h3.find_next('ul'):定位当前H3标签之后的第一个UL标签,确保是该标题对应的列表
  • get_text(strip=True):提取标签内的文本,并自动去除前后多余的空格和换行符

运行代码后,输出格式就是每个标题搭配其下所有LI的文本内容,完全符合你的需求。

内容的提问来源于stack exchange,提问作者Shubham Kapadnis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 16:15:35