如何从div内多个<p>元素提取单个数据?代码输出空行求助
Hey there! Let's figure out why you're getting empty output when trying to extract "Test1" from the details-organisation div. Here are the most common issues and actionable fixes:
I'll assume you're using a common tool like BeautifulSoup (the go-to for static HTML parsing) first, then cover dynamic content scenarios too.
1. Your Selector Isn't Matching the Div Correctly
A frequent mistake is mis-targeting the details-organisation div—either due to syntax errors or ignoring Python's keyword restrictions.
Example Correct Selector (BeautifulSoup):
If your HTML structure looks like this:
<div class="details-organisation"> <p>Test1</p> <p>Other irrelevant text</p> </div>
Use class_ (with an underscore) instead of class (since class is a reserved Python keyword):
from bs4 import BeautifulSoup # Your HTML content (loaded from a file/request or hardcoded) html = ''' <div class="details-organisation"> <p>Test1</p> <p>Other irrelevant text</p> </div> ''' soup = BeautifulSoup(html, 'html.parser') # Correctly target the div org_div = soup.find('div', class_='details-organisation') # First, check if the div was found to avoid errors if not org_div: print("Error: Could not locate the 'details-organisation' div") else: # Extract text from the first <p> tag, stripping extra whitespace target_text = org_div.find('p').get_text(strip=True) print(target_text) # Outputs: Test1
2. You're Extracting Text the Wrong Way
Calling .text directly on the div will return combined text from all child <p> elements. To isolate "Test1", target the specific <p> tag (usually the first one) and use get_text(strip=True) to eliminate hidden newlines/spaces that might make output appear empty.
3. The Content is Dynamically Loaded via JavaScript
If the details-organisation div only appears after the page runs JavaScript (e.g., loaded from an API), static parsers like BeautifulSoup can't see it. Use a browser automation tool like Selenium to render the full page:
Selenium Example:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize Chrome driver (ensure ChromeDriver is installed) driver = webdriver.Chrome() driver.get("https://your-target-page-url.com") # Wait up to 10 seconds for the div to load try: org_div = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "details-organisation")) ) # Extract cleaned text from the first <p> tag target_text = org_div.find_element(By.TAG_NAME, "p").text.strip() print(target_text) # Outputs: Test1 finally: driver.quit()
4. Hidden Whitespace is Causing "Empty" Output
Sometimes the <p> tag contains excessive newlines or spaces that make output look blank. Using strip=True (for BeautifulSoup) or .strip() (for Selenium) will clean up this whitespace and reveal the actual text.
内容的提问来源于stack exchange,提问作者BIGJOHN

