如何用Beautiful Soup提取HTML数据并存入列表及解决常见问题
Hey there! Let's tackle your Beautiful Soup questions one by one, and fix those frustrating errors you're hitting.
src Attribute from the img Tag Target HTML:
<div id="headshot"> <img title="Photo of someone" alt="Photo of somenone" src="url/file.jpg"> </div>Solution:
Use the uniqueid="headshot"to target the parent div first, then grab thesrcattribute from the nestedimgtag. Adding checks prevents crashes if the HTML is malformed:# Method 1: Chain find() calls headshot_div = soup.find('div', id='headshot') if headshot_div: img_src = headshot_div.find('img')['src'] print(img_src) # Output: url/file.jpg # Method 2: Concise CSS selector img_src = soup.select_one('#headshot img')['src']
bioheading and biodata Values into a List Target HTML:
<p class="bioheading">value</p> <div class="biodata">value</div> <p class="bioheading">value</p> <p class="biodata">value</p> ...Solution:
Grab all elements with either class, then pair them into a list of tuples (or a flat list if you just need all values):# Get all relevant elements (matches p/div tags with target classes) bio_elements = soup.find_all(['p', 'div'], class_=['bioheading', 'biodata']) # Option 1: Store as (heading, data) pairs bio_pairs = [] for i in range(0, len(bio_elements), 2): if i + 1 < len(bio_elements): # Avoid index out of bounds heading = bio_elements[i].get_text(strip=True) data = bio_elements[i+1].get_text(strip=True) bio_pairs.append((heading, data)) # Option 2: Flat list of all values flat_bio_values = [elem.get_text(strip=True) for elem in bio_elements]
strip=Truecleans up extra whitespace from messy HTML.
bioheading and biodata Elements Target HTML:
<div id="vitalbox" class="tab-content"> <div role="tabpanel" class="tab-pane active" id="home"> <div class="row"> <div class="col-xs-12 col-sm-4"> <p class="bioheading">value</p> ... </div>Solution:
First locate the innermost container with a CSS selector, then extract the elements just like in question 2:# Target the specific nested container bio_container = soup.select_one('#vitalbox .tab-pane.active .col-xs-12.col-sm-4') if bio_container: bio_elements = bio_container.find_all('p', class_=['bioheading', 'biodata']) # Reuse the pairing logic from question 2 bio_pairs = [] for i in range(0, len(bio_elements), 2): if i + 1 < len(bio_elements): heading = bio_elements[i].get_text(strip=True) data = bio_elements[i+1].get_text(strip=True) bio_pairs.append((heading, data))The CSS selector ensures you only pull data from the active tab panel, ignoring hidden content.
Target HTML:
<input name="Gender" value="m" type="hidden">Solution:
Your original code crashed because some<input>tags don't have anameattribute. Use.get()to safely check for attributes instead of accessing them directly:# Narrow down to hidden inputs first to reduce loops for hidden_input in soup.find_all('input', type='hidden'): # Check if 'name' exists and equals 'Gender' if hidden_input.get('name') == 'Gender': gender_val = hidden_input.get('value', '') if gender_val == 'f': out.write('F') elif gender_val == 'm': out.write('M') else: out.write('—') break # Stop looping once we find the Gender input
.get('attribute')returnsNoneif the attribute doesn't exist, so no more KeyError!
Error 1: TypeError: 'NoneType' object is not callable with .find_All()
This is a simple typo! Beautiful Soup's method is find_all() (all lowercase), not find_All() (mixed case). Your switch to findAll() (old-style syntax) works, but find_all() is the official recommended method. Here's your fixed link extraction code:
import re k = 0 a_table = [] bday1 = '' for link in soup.find_all('a'): href = link.get('href') a_table.append(str(href) if href else '') # Handle None href values if href and re.match(regs4, href, re.M): bday_match = re.search(regs1, href, re.M) if bday_match: bday1 = bday_match.group() # Don't forget to get the matched string! else: bday1 = 'http://url.com/calendar.asp?calmonth=01&calyear=2018&calday=01' else: bday1 = 'http://url.com/calendar.asp?calmonth=01&calyear=2018&calday=01' k += 1
- Added checks for
Nonehref values to avoid passing invalid data to regex. - Used
.group()to get the actual matched string fromre.search()(your original code stored the match object instead of the value).
Error 2: KeyError: 'name' in Gender Extraction
As covered in question 4, replacing direct attribute access with .get() fixes this. The key is to never assume an attribute exists in messy HTML!
内容的提问来源于stack exchange,提问作者RyosanCiffer

