You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup提取HTML数据并存入列表及解决常见问题

Hey there! Let's tackle your Beautiful Soup questions one by one, and fix those frustrating errors you're hitting.


1. Extract src Attribute from the img Tag

Target HTML:

<div id="headshot"> <img title="Photo of someone" alt="Photo of somenone" src="url/file.jpg"> </div>

Solution:
Use the unique id="headshot" to target the parent div first, then grab the src attribute from the nested img tag. Adding checks prevents crashes if the HTML is malformed:

# Method 1: Chain find() calls
headshot_div = soup.find('div', id='headshot')
if headshot_div:
    img_src = headshot_div.find('img')['src']
    print(img_src)  # Output: url/file.jpg

# Method 2: Concise CSS selector
img_src = soup.select_one('#headshot img')['src']
2. Extract bioheading and biodata Values into a List

Target HTML:

<p class="bioheading">value</p> <div class="biodata">value</div> <p class="bioheading">value</p> <p class="biodata">value</p> ...

Solution:
Grab all elements with either class, then pair them into a list of tuples (or a flat list if you just need all values):

# Get all relevant elements (matches p/div tags with target classes)
bio_elements = soup.find_all(['p', 'div'], class_=['bioheading', 'biodata'])

# Option 1: Store as (heading, data) pairs
bio_pairs = []
for i in range(0, len(bio_elements), 2):
    if i + 1 < len(bio_elements):  # Avoid index out of bounds
        heading = bio_elements[i].get_text(strip=True)
        data = bio_elements[i+1].get_text(strip=True)
        bio_pairs.append((heading, data))

# Option 2: Flat list of all values
flat_bio_values = [elem.get_text(strip=True) for elem in bio_elements]

strip=True cleans up extra whitespace from messy HTML.

3. Extract Nested bioheading and biodata Elements

Target HTML:

<div id="vitalbox" class="tab-content"> <div role="tabpanel" class="tab-pane active" id="home"> <div class="row"> <div class="col-xs-12 col-sm-4"> <p class="bioheading">value</p> ... </div>

Solution:
First locate the innermost container with a CSS selector, then extract the elements just like in question 2:

# Target the specific nested container
bio_container = soup.select_one('#vitalbox .tab-pane.active .col-xs-12.col-sm-4')

if bio_container:
    bio_elements = bio_container.find_all('p', class_=['bioheading', 'biodata'])
    # Reuse the pairing logic from question 2
    bio_pairs = []
    for i in range(0, len(bio_elements), 2):
        if i + 1 < len(bio_elements):
            heading = bio_elements[i].get_text(strip=True)
            data = bio_elements[i+1].get_text(strip=True)
            bio_pairs.append((heading, data))

The CSS selector ensures you only pull data from the active tab panel, ignoring hidden content.

4. Extract Hidden Input's Gender Value (Fixing KeyError)

Target HTML:

<input name="Gender" value="m" type="hidden">

Solution:
Your original code crashed because some <input> tags don't have a name attribute. Use .get() to safely check for attributes instead of accessing them directly:

# Narrow down to hidden inputs first to reduce loops
for hidden_input in soup.find_all('input', type='hidden'):
    # Check if 'name' exists and equals 'Gender'
    if hidden_input.get('name') == 'Gender':
        gender_val = hidden_input.get('value', '')
        if gender_val == 'f':
            out.write('F')
        elif gender_val == 'm':
            out.write('M')
        else:
            out.write('—')
        break  # Stop looping once we find the Gender input

.get('attribute') returns None if the attribute doesn't exist, so no more KeyError!


Fixing Your Existing Code Errors

Error 1: TypeError: 'NoneType' object is not callable with .find_All()

This is a simple typo! Beautiful Soup's method is find_all() (all lowercase), not find_All() (mixed case). Your switch to findAll() (old-style syntax) works, but find_all() is the official recommended method. Here's your fixed link extraction code:

import re

k = 0
a_table = []
bday1 = ''
for link in soup.find_all('a'):
    href = link.get('href')
    a_table.append(str(href) if href else '')  # Handle None href values
    if href and re.match(regs4, href, re.M):
        bday_match = re.search(regs1, href, re.M)
        if bday_match:
            bday1 = bday_match.group()  # Don't forget to get the matched string!
        else:
            bday1 = 'http://url.com/calendar.asp?calmonth=01&amp;calyear=2018&amp;calday=01'
    else:
        bday1 = 'http://url.com/calendar.asp?calmonth=01&amp;calyear=2018&amp;calday=01'
    k += 1
  • Added checks for None href values to avoid passing invalid data to regex.
  • Used .group() to get the actual matched string from re.search() (your original code stored the match object instead of the value).

Error 2: KeyError: 'name' in Gender Extraction

As covered in question 4, replacing direct attribute access with .get() fixes this. The key is to never assume an attribute exists in messy HTML!


内容的提问来源于stack exchange,提问作者RyosanCiffer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:27:02