You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup根据前置标签内容获取指定div的内容?

根据前置H4标签获取对应profile-area内容的解决方案

看起来你已经能抓取到所有profile-area元素的文本了,但现在需要把它们和前面的h4标题精准对应起来对吧?我来帮你调整代码,既能关联单个标题对应单个内容的情况,也能处理像Output这种一个标题对应多个内容的场景。

首先先修正你代码里的一个小问题:你定义的HTML变量是html_doc,但初始化BeautifulSoup的时候用了html,这个要改成一致的,不然会报错。

接下来是核心解决方案:我们可以遍历每个h4标签,然后逐个获取它后面连续的profile-area元素,直到遇到非目标元素为止。这样就能完美对应每个标题下的所有内容了。

完整代码示例:

import requests
from bs4 import BeautifulSoup

html_doc = """
<html>
<body>
<div class="col-md-6">
<iframe class="factory_detail_google_map" frameborder="0" src= "https://www.google.com/maps/embed/v1/search?q=3.037787%2C101.38189&amp;key=AIzaSyCMDADp9QHYbQ8OBGl8puAOv-16W8ziz7Y" allowfullscreen=""></iframe>
</div>
<div class="col-md-12">
<h4>Models &amp; Products</h4>
<div class="profile-area"> Large Buses, Trucks, Trailer-heads </div>
<h4>Production Capacity (year)</h4>
<div class="profile-area"> Vehicle 700 units /year </div>
<h4>Output</h4>
<div class="profile-area"> Vehicle 356 units ( 2016 ) </div>
<div class="profile-area"> Vehicle 477 units ( 2015 ) </div>
<div class="profile-area"> Vehicle 760 units ( 2014 ) </div>
<div class="profile-area"> Vehicle 647 units ( 2013 ) </div>
</div>
</body>
</html>
"""

# 修正变量名,用html_doc代替原代码中的html
soup = BeautifulSoup(html_doc, 'lxml')

# 遍历col-md-12容器下的所有h4标签,避免抓取无关内容
for section_title in soup.select("div.col-md-12 h4"):
    # 清理标题文本的多余空格,让输出更整洁
    title_text = section_title.get_text(strip=True)
    print(f"=== {title_text} ===")
    
    # 从h4的下一个兄弟元素开始依次查找
    current_element = section_title.next_sibling
    while current_element:
        # 检查当前元素是否是目标的profile-area类div
        if current_element.name == "div" and "profile-area" in current_element.get("class", []):
            # 输出清理后的内容
            print(current_element.get_text(strip=True))
            # 继续查找下一个兄弟元素
            current_element = current_element.next_sibling
        else:
            # 遇到非目标元素,停止当前标题的内容查找
            break

# 额外补充你提到的谷歌地图坐标拆分参考(你说自己能解决,这里给个简洁实现)
map_iframe = soup.select_one("iframe.factory_detail_google_map")
if map_iframe:
    map_src = map_iframe.get("src")
    # 提取q参数中的坐标部分并拆分
    coord_part = map_src.split("q=")[1].split("&")[0]
    latitude, longitude = coord_part.split("%2C")
    print(f"\n谷歌地图坐标:纬度 {latitude},经度 {longitude}")

代码逻辑说明:

  • 先锁定col-md-12容器下的h4标签,避免抓取页面其他区域的无关标题
  • 对每个标题,逐个检查后续兄弟元素:如果是profile-area类的div就输出内容,否则跳出循环处理下一个标题
  • 用get_text(strip=True)自动清理文本前后的空格,让输出更规整

这样运行代码后,你就能得到每个标题对应的所有内容,完全匹配你的需求啦。

内容的提问来源于stack exchange,提问作者Luis Paganini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:01:16