网页爬取技术咨询:如何从指定网站单独提取<p>标签的部分数据而非一次性获取所有标签数据
Hey there! Let's tackle this problem together—getting only the specific
content you need from that Animal Diversity Web page instead of pulling every single paragraph. Since you didn’t share your exact code, I’ll walk you through common, reliable approaches using Python’s BeautifulSoup (the standard tool for web scraping like this).
Start by identifying exactly where your desired content lives in the page’s HTML:
- Right-click on the specific text you want to extract
- Select Inspect (in Chrome, Firefox, or Edge)
- Look at the surrounding code: Do the
tags sit inside a specific div with a class (like
attribute-data, which is used heavily on that site)? Are they directly after a heading (like<h3>Geographic Range</h3>)? This context is your key to filtering.
Based on the page you linked, here are tailored ways to narrow down your results:
Scenario 1: Target
Tags Inside a Specific Parent Container
Most content on that site is grouped inside divs with the class attribute-data. If your desired paragraphs are in one of these, you can first grab the container, then pull only its
tags:
import requests from bs4 import BeautifulSoup url = "https://animaldiversity.org/accounts/Callithrix%20humilis" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # Locate the parent container (adjust the class name to match your target section) target_container = soup.find('div', class_='attribute-data') # Extract only <p> tags within this container target_paragraphs = target_container.find_all('p') # Loop through and get clean text for p in target_paragraphs: print(p.get_text(strip=True))
Scenario 2: Target
Tags Immediately After a Specific Heading
If your desired content follows a clear heading (like "Conservation Status"), use find_next_sibling to grab the exact paragraph right after it:
# Find the relevant heading first target_heading = soup.find('h3', string="Conservation Status") # Get the next <p> tag directly after the heading target_p = target_heading.find_next_sibling('p') if target_p: print(target_p.get_text(strip=True))
Scenario 3: Filter
Tags by Keyword Content
If you only want paragraphs that mention specific terms (like "diet" or "habitat"), filter all
tags with a simple list comprehension:
all_paragraphs = soup.find_all('p') # Keep only paragraphs containing your keyword (case-insensitive) target_paragraphs = [p for p in all_paragraphs if "diet" in p.get_text(strip=True).lower()] for p in target_paragraphs: print(p.get_text(strip=True))
Always rely on the page’s structured containers (like those attribute-data divs) instead of grabbing every
tag. This makes your code more resilient if the site makes small layout changes later.
内容的提问来源于stack exchange,提问作者saurav bhatt

