You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup4问题:通过strong标签值识别信息仅对部分标签值生效

Extracting Dynamic Data Points with BeautifulSoup

Hey there! Let's figure out how to grab those tricky "Accepts Business From:" and "Classes of business" values, even when the page's class names change and elements are out of order. The key here is to target the label text instead of unreliable class names—since the labels themselves are consistent, that's our solid anchor.

Extract "Accepts Business From:" Value

We'll search for the element containing the target label text, then traverse the HTML to pull its corresponding value. Here's a flexible approach that works regardless of element order:

# Initialize variable to store the extracted value
accepts_business_from = None

# Find the element with our label text (handles extra whitespace variations)
label_element = company.find(string=lambda text: text and "Accepts Business From:" in text.strip())

if label_element:
    # Adjust this traversal to match your actual HTML structure
    # Common scenarios:
    # 1. Value lives in the next sibling tag
    value_tag = label_element.parent.find_next_sibling()
    # 2. Value is a child of the same parent (e.g., <div><label>...</label><span>VALUE</span></div>)
    # value_tag = label_element.parent.find("span")
    
    if value_tag:
        accepts_business_from = value_tag.get_text(strip=True)

Extract "Classes of business" Value

This field might be a single value or a list—either way, we use the same label-targeting method to fetch it reliably:

# Initialize variable to store the value(s)
classes_of_business = None

# Locate the label element
label_element = company.find(string=lambda text: text and "Classes of business" in text.strip())

if label_element:
    # If the value is a list (e.g., <ul> with <li> items)
    list_container = label_element.parent.find_next_sibling("ul")
    if list_container:
        classes_of_business = [item.get_text(strip=True) for item in list_container.find_all("li")]
    else:
        # If it's a single value, use the same sibling/child traversal logic as before
        value_tag = label_element.parent.find_next_sibling()
        classes_of_business = value_tag.get_text(strip=True) if value_tag else None

Quick Tips

  • Inspect your target HTML to adjust the traversal (like find_next_sibling() or find())—this ensures you're grabbing the right element relative to the label.
  • The lambda in find(string=...) handles extra spaces or line breaks around the label text, making the match more robust than exact string comparisons.

内容的提问来源于stack exchange,提问作者Richard Golz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:25:42