使用BeautifulSoup4获取非唯一标签文本及Classes of business数据问题
Hey there! Since you're already familiar with BeautifulSoup and Python, let's jump right into solving this issue. The problem with your strong.text.strip() == 'Classes of business' check almost always boils down to hidden whitespace characters (like non-breaking spaces or newlines) that strip() isn't handling properly. Let's walk through a reliable way to loop through each company entry and grab the data you need.
First: Why Your Check Isn't Working
Most often, the <strong> tag's text isn't exactly "Classes of business" when you look under the hood. Common culprits include:
- Non-breaking spaces (
in HTML, which becomes\xa0in Python) thatstrip()doesn't remove by default. - Hidden newlines or tabs inside the
<strong>tag (e.g., the text is split across lines in the HTML). - Minor typos or extra spaces in the label (like "Classes of business " with a trailing space you can't see).
The fix here is to use get_text(strip=True) instead of text.strip()—this method handles all types of whitespace (including Unicode spaces) and trims them from both ends.
Step-by-Step Traversal Method
Assuming your HTML structure for each company looks something like this (based on typical directory layouts):
<!-- Company 1 --> <div class="company-record"> <div class="detail-row"> <strong>Company Name:</strong> <span>Global Insurance Group</span> </div> <div class="detail-row"> <strong>Classes of business:</strong> <span>Property, Casualty, Marine</span> </div> </div> <!-- Company 2 --> <div class="company-record"> <div class="detail-row"> <strong>Company Name:</strong> <span>LifeSecure Providers</span> </div> <div class="detail-row"> <strong>Classes of business:</strong> <span>Life, Health, Disability</span> </div> </div>
Here's how to scrape the "Classes of business" values for all 10 entries:
from bs4 import BeautifulSoup # Assume you've already loaded your HTML into 'soup' soup = BeautifulSoup(your_html_content, 'html.parser') # Initialize a list to store the values classes_list = [] # First, grab all individual company records (adjust the selector to match your actual HTML) company_records = soup.find_all('div', class_='company-record') # Loop through each company record for record in company_records: # Look through all <strong> tags inside the current company's block for strong_tag in record.find_all('strong'): # Get the cleaned label text label = strong_tag.get_text(strip=True) # Check if it's exactly the label we want if label == 'Classes of business': # Grab the corresponding value—adjust this selector to match your HTML structure # This example assumes the value is in the next sibling <span> business_class = strong_tag.find_next_sibling('span').get_text(strip=True) classes_list.append(business_class) # Break out of the inner loop since we found what we need for this company break # Print the result to verify print(classes_list)
Adjustments for Your Specific HTML
If your value isn't in a <span> next to the <strong> tag, here are alternative ways to grab it:
- If the value is in the same parent div but after the strong tag:
strong_tag.parent.find(text=True, recursive=False).strip() - If there's a specific class on the value element:
strong_tag.find_next('div', class_='value').get_text(strip=True)
Troubleshooting Tip
If the label still doesn't match, print out the raw text and its repr to see hidden characters:
for strong_tag in soup.find_all('strong'): print(repr(strong_tag.get_text()))
This will show you exactly what's in the text (like 'Classes of business\xa0' if there's a non-breaking space), so you can adjust your check accordingly (e.g., label.replace('\xa0', '').strip() == 'Classes of business').
内容的提问来源于stack exchange,提问作者Richard Golz

