You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取技术咨询:如何从指定网站单独提取<p>标签的部分数据而非一次性获取所有标签数据

Hey there! Let's tackle this problem together—getting only the specific

content you need from that Animal Diversity Web page instead of pulling every single paragraph. Since you didn’t share your exact code, I’ll walk you through common, reliable approaches using Python’s BeautifulSoup (the standard tool for web scraping like this).

1. First, Pinpoint Your Target

Tags

Start by identifying exactly where your desired content lives in the page’s HTML:

  • Right-click on the specific text you want to extract
  • Select Inspect (in Chrome, Firefox, or Edge)
  • Look at the surrounding code: Do the

    tags sit inside a specific div with a class (like attribute-data, which is used heavily on that site)? Are they directly after a heading (like <h3>Geographic Range</h3>)? This context is your key to filtering.

2. Example Solutions for Common Scenarios

Based on the page you linked, here are tailored ways to narrow down your results:

Scenario 1: Target

Tags Inside a Specific Parent Container

Most content on that site is grouped inside divs with the class attribute-data. If your desired paragraphs are in one of these, you can first grab the container, then pull only its

tags:

import requests
from bs4 import BeautifulSoup

url = "https://animaldiversity.org/accounts/Callithrix%20humilis"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# Locate the parent container (adjust the class name to match your target section)
target_container = soup.find('div', class_='attribute-data')

# Extract only <p> tags within this container
target_paragraphs = target_container.find_all('p')

# Loop through and get clean text
for p in target_paragraphs:
    print(p.get_text(strip=True))

Scenario 2: Target

Tags Immediately After a Specific Heading

If your desired content follows a clear heading (like "Conservation Status"), use find_next_sibling to grab the exact paragraph right after it:

# Find the relevant heading first
target_heading = soup.find('h3', string="Conservation Status")

# Get the next <p> tag directly after the heading
target_p = target_heading.find_next_sibling('p')

if target_p:
    print(target_p.get_text(strip=True))

Scenario 3: Filter

Tags by Keyword Content

If you only want paragraphs that mention specific terms (like "diet" or "habitat"), filter all

tags with a simple list comprehension:

all_paragraphs = soup.find_all('p')
# Keep only paragraphs containing your keyword (case-insensitive)
target_paragraphs = [p for p in all_paragraphs if "diet" in p.get_text(strip=True).lower()]

for p in target_paragraphs:
    print(p.get_text(strip=True))
3. Quick Pro Tip

Always rely on the page’s structured containers (like those attribute-data divs) instead of grabbing every

tag. This makes your code more resilient if the site makes small layout changes later.

内容的提问来源于stack exchange,提问作者saurav bhatt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 21:57:48