使用Beautiful Soup提取H2标签中的文本内容
Fixing Your Beautiful Soup Text Extraction Code
Let's break down what's going wrong with your current code, then fix it to extract the text from those <h2> tags correctly.
What's Off in Your Current Code
- Each
itemin yourdatalist is already an<h2>tag object. When you callitem.findAll('h2'), you're asking that single h2 tag to find more h2 tags inside itself (which don't exist), and this returns a list of elements. Lists don't have aget_text()method—that's why your code throws an error. - Your code also has indentation issues: Python relies on indentation to define code blocks, so the
forloop needs to be properly indented under the variable assignment.
Correct Approach
Assuming you've already parsed your HTML with Beautiful Soup, here's how to extract the text properly:
Step 1: Parse HTML and Fetch Target H2 Elements
First, correctly grab all the <h2> tags using their specific class:
from bs4 import BeautifulSoup # Replace this with your actual HTML content html_content = """ <h2 class="ArticleBody-styles-makeit-subtitle--LnHeO">1. Python</h2> <h2 class="ArticleBody-styles-makeit-subtitle--LnHeO">2. HTML</h2> """ # Parse the HTML soup = BeautifulSoup(html_content, "html.parser") # Get all h2 tags matching the class data = soup.find_all("h2", class_="ArticleBody-styles-makeit-subtitle--LnHeO")
Step 2: Extract Text from Each Tag
Now that data is a list of <h2> tag objects, loop through them and call get_text() directly on each item:
for item in data: # Extract and print the inner text print(item.get_text())
This will output exactly what you need:
1. Python 2. HTML
Key Notes
- Use
class_instead ofclassin thefind_all()call—classis a reserved keyword in Python, so Beautiful Soup usesclass_to target CSS classes. - If you only need a single
<h2>tag instead of all matching ones, usesoup.find()(singular) instead offind_all()—this returns a single tag object, not a list.
内容的提问来源于stack exchange,提问作者Raja G
相关产品推荐
相关产品推荐

