BeautifulSoup4问题:通过strong标签值识别信息仅对部分标签值生效
Extracting Dynamic Data Points with BeautifulSoup
Hey there! Let's figure out how to grab those tricky "Accepts Business From:" and "Classes of business" values, even when the page's class names change and elements are out of order. The key here is to target the label text instead of unreliable class names—since the labels themselves are consistent, that's our solid anchor.
Extract "Accepts Business From:" Value
We'll search for the element containing the target label text, then traverse the HTML to pull its corresponding value. Here's a flexible approach that works regardless of element order:
# Initialize variable to store the extracted value accepts_business_from = None # Find the element with our label text (handles extra whitespace variations) label_element = company.find(string=lambda text: text and "Accepts Business From:" in text.strip()) if label_element: # Adjust this traversal to match your actual HTML structure # Common scenarios: # 1. Value lives in the next sibling tag value_tag = label_element.parent.find_next_sibling() # 2. Value is a child of the same parent (e.g., <div><label>...</label><span>VALUE</span></div>) # value_tag = label_element.parent.find("span") if value_tag: accepts_business_from = value_tag.get_text(strip=True)
Extract "Classes of business" Value
This field might be a single value or a list—either way, we use the same label-targeting method to fetch it reliably:
# Initialize variable to store the value(s) classes_of_business = None # Locate the label element label_element = company.find(string=lambda text: text and "Classes of business" in text.strip()) if label_element: # If the value is a list (e.g., <ul> with <li> items) list_container = label_element.parent.find_next_sibling("ul") if list_container: classes_of_business = [item.get_text(strip=True) for item in list_container.find_all("li")] else: # If it's a single value, use the same sibling/child traversal logic as before value_tag = label_element.parent.find_next_sibling() classes_of_business = value_tag.get_text(strip=True) if value_tag else None
Quick Tips
- Inspect your target HTML to adjust the traversal (like
find_next_sibling()orfind())—this ensures you're grabbing the right element relative to the label. - The
lambdainfind(string=...)handles extra spaces or line breaks around the label text, making the match more robust than exact string comparisons.
内容的提问来源于stack exchange,提问作者Richard Golz
相关产品推荐
相关产品推荐

