如何用BeautifulSoup提取<li>内<a>标签的指定文本?
<a> Tag Text Hey there! Let's sort out why li.a.string isn't giving you the "this word" text you're after, and go through reliable fixes.
Why li.a.string might fail
The .string attribute in BeautifulSoup only returns a text value if the tag contains exactly one text node with no other child elements. Even if there's hidden whitespace or line breaks around your target text, .string might return None or unexpected results. Instead, use these more robust methods:
Method 1: Use .get_text(strip=True) (Most Recommended)
This method extracts all text inside the tag, and the strip=True parameter cleans up any extra whitespace, newlines, or tabs around your target text:
result = li.a.get_text(strip=True)
Method 2: Access the first content node
If you're sure the text is the first child of the <a> tag, you can directly grab it with .contents and clean up whitespace:
result = li.a.contents[0].strip()
Method 3: Use the .text attribute
Similar to .get_text(), the .text attribute pulls all text from the tag. Just remember to strip extra whitespace:
result = li.a.text.strip()
Full Example Code
Here's how it all comes together with your HTML snippet:
from bs4 import BeautifulSoup # Your target HTML html = '<li> <a href="/3137?category=957259">"this word"</a> <span class="date">11:08:57</span> </li>' soup = BeautifulSoup(html, 'html.parser') li = soup.find('li') # Test all methods print(li.a.get_text(strip=True)) # Output: "this word" print(li.a.contents[0].strip()) # Output: "this word" print(li.a.text.strip()) # Output: "this word"
All these approaches will reliably get you the "this word" text you need. .get_text(strip=True) is the most versatile option since it handles edge cases like nested tags or extra whitespace automatically.
内容的提问来源于stack exchange,提问作者Sang_Hyeop

