You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从div内多个<p>元素提取单个数据?代码输出空行求助

Hey there! Let's figure out why you're getting empty output when trying to extract "Test1" from the details-organisation div. Here are the most common issues and actionable fixes:

Troubleshooting Empty Output When Extracting "Test1" from Your Div

I'll assume you're using a common tool like BeautifulSoup (the go-to for static HTML parsing) first, then cover dynamic content scenarios too.

1. Your Selector Isn't Matching the Div Correctly

A frequent mistake is mis-targeting the details-organisation div—either due to syntax errors or ignoring Python's keyword restrictions.

Example Correct Selector (BeautifulSoup):

If your HTML structure looks like this:

<div class="details-organisation">
  <p>Test1</p>
  <p>Other irrelevant text</p>
</div>

Use class_ (with an underscore) instead of class (since class is a reserved Python keyword):

from bs4 import BeautifulSoup

# Your HTML content (loaded from a file/request or hardcoded)
html = '''
<div class="details-organisation">
  <p>Test1</p>
  <p>Other irrelevant text</p>
</div>
'''

soup = BeautifulSoup(html, 'html.parser')
# Correctly target the div
org_div = soup.find('div', class_='details-organisation')

# First, check if the div was found to avoid errors
if not org_div:
    print("Error: Could not locate the 'details-organisation' div")
else:
    # Extract text from the first <p> tag, stripping extra whitespace
    target_text = org_div.find('p').get_text(strip=True)
    print(target_text)  # Outputs: Test1

2. You're Extracting Text the Wrong Way

Calling .text directly on the div will return combined text from all child <p> elements. To isolate "Test1", target the specific <p> tag (usually the first one) and use get_text(strip=True) to eliminate hidden newlines/spaces that might make output appear empty.

3. The Content is Dynamically Loaded via JavaScript

If the details-organisation div only appears after the page runs JavaScript (e.g., loaded from an API), static parsers like BeautifulSoup can't see it. Use a browser automation tool like Selenium to render the full page:

Selenium Example:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize Chrome driver (ensure ChromeDriver is installed)
driver = webdriver.Chrome()
driver.get("https://your-target-page-url.com")

# Wait up to 10 seconds for the div to load
try:
    org_div = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "details-organisation"))
    )
    # Extract cleaned text from the first <p> tag
    target_text = org_div.find_element(By.TAG_NAME, "p").text.strip()
    print(target_text)  # Outputs: Test1
finally:
    driver.quit()

4. Hidden Whitespace is Causing "Empty" Output

Sometimes the <p> tag contains excessive newlines or spaces that make output look blank. Using strip=True (for BeautifulSoup) or .strip() (for Selenium) will clean up this whitespace and reveal the actual text.


内容的提问来源于stack exchange,提问作者BIGJOHN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:48:15