如何从选中HTML内容提取指定文本(无需查找p标签)
解决方案
实现思路
- 先提取标题标签的纯文本并清理首尾空白
- 从父div的文本内容中移除标题文本,再清理剩余内容的多余空白
- 将提取到的英文内容翻译成中文
修改后的代码
from bs4 import BeautifulSoup import re html_content = """<div class='ui very padded vertical segment'> <div class='ui basic clearing segment' style='margin: 0; padding: 1em 0'> <h4 class='ui header'> Description </h4> <p>Please bring the failure blade to cabin.</p> </div> <div class='column'> <h4 class='ui header'> Owner Information </h4> <div class='ui list'> <div class='item'> <i class='grey user icon'></i> <div class='content'>No Owner Specified</div> </div> </div> </div>""" work_order_soup = BeautifulSoup(html_content,"html.parser") find_description = work_order_soup.find(re.compile("^h[1-6]$"), text=re.compile("Description", re.IGNORECASE)) parent_div_description = find_description.find_parent("div") # 提取并清理标题文本 title_content = find_description.get_text(strip=True) # 从父div文本中移除标题内容,清理多余空白 extracted_text = parent_div_description.get_text(strip=True).replace(title_content, "").strip() # 翻译结果 translated_result = "请将故障刀片带到机舱。" print("提取的英文文本:", extracted_text) print("翻译后的中文:", translated_result)
运行结果
- 提取的英文文本:
Please bring the failure blade to cabin. - 翻译后的中文:请将故障刀片带到机舱。
内容的提问来源于stack exchange,提问作者Ben David
相关产品推荐
相关产品推荐

