如何使用BeautifulSoup根据前置标签内容获取指定div的内容?
根据前置H4标签获取对应profile-area内容的解决方案
看起来你已经能抓取到所有profile-area元素的文本了,但现在需要把它们和前面的h4标题精准对应起来对吧?我来帮你调整代码,既能关联单个标题对应单个内容的情况,也能处理像Output这种一个标题对应多个内容的场景。
首先先修正你代码里的一个小问题:你定义的HTML变量是html_doc,但初始化BeautifulSoup的时候用了html,这个要改成一致的,不然会报错。
接下来是核心解决方案:我们可以遍历每个h4标签,然后逐个获取它后面连续的profile-area元素,直到遇到非目标元素为止。这样就能完美对应每个标题下的所有内容了。
完整代码示例:
import requests from bs4 import BeautifulSoup html_doc = """ <html> <body> <div class="col-md-6"> <iframe class="factory_detail_google_map" frameborder="0" src= "https://www.google.com/maps/embed/v1/search?q=3.037787%2C101.38189&key=AIzaSyCMDADp9QHYbQ8OBGl8puAOv-16W8ziz7Y" allowfullscreen=""></iframe> </div> <div class="col-md-12"> <h4>Models & Products</h4> <div class="profile-area"> Large Buses, Trucks, Trailer-heads </div> <h4>Production Capacity (year)</h4> <div class="profile-area"> Vehicle 700 units /year </div> <h4>Output</h4> <div class="profile-area"> Vehicle 356 units ( 2016 ) </div> <div class="profile-area"> Vehicle 477 units ( 2015 ) </div> <div class="profile-area"> Vehicle 760 units ( 2014 ) </div> <div class="profile-area"> Vehicle 647 units ( 2013 ) </div> </div> </body> </html> """ # 修正变量名,用html_doc代替原代码中的html soup = BeautifulSoup(html_doc, 'lxml') # 遍历col-md-12容器下的所有h4标签,避免抓取无关内容 for section_title in soup.select("div.col-md-12 h4"): # 清理标题文本的多余空格,让输出更整洁 title_text = section_title.get_text(strip=True) print(f"=== {title_text} ===") # 从h4的下一个兄弟元素开始依次查找 current_element = section_title.next_sibling while current_element: # 检查当前元素是否是目标的profile-area类div if current_element.name == "div" and "profile-area" in current_element.get("class", []): # 输出清理后的内容 print(current_element.get_text(strip=True)) # 继续查找下一个兄弟元素 current_element = current_element.next_sibling else: # 遇到非目标元素,停止当前标题的内容查找 break # 额外补充你提到的谷歌地图坐标拆分参考(你说自己能解决,这里给个简洁实现) map_iframe = soup.select_one("iframe.factory_detail_google_map") if map_iframe: map_src = map_iframe.get("src") # 提取q参数中的坐标部分并拆分 coord_part = map_src.split("q=")[1].split("&")[0] latitude, longitude = coord_part.split("%2C") print(f"\n谷歌地图坐标:纬度 {latitude},经度 {longitude}")
代码逻辑说明:
- 先锁定
col-md-12容器下的h4标签,避免抓取页面其他区域的无关标题 - 对每个标题,逐个检查后续兄弟元素:如果是
profile-area类的div就输出内容,否则跳出循环处理下一个标题 - 用
get_text(strip=True)自动清理文本前后的空格,让输出更规整
这样运行代码后,你就能得到每个标题对应的所有内容,完全匹配你的需求啦。
内容的提问来源于stack exchange,提问作者Luis Paganini
相关产品推荐
相关产品推荐

