如何在BeautifulSoup中用正则提取<p>标签文本?Python脚本返回空列表
<p> Tag with Regex Hey Rajat, let's break down why your Python regex is returning an empty list and get you the text you need from that <p> tag.
Common Regex Pitfalls Here
The most likely issues with your original regex are:
- Not accounting for multiple class values: Your target
<p>has two classes (result__outcomeandu-hide-phablet), so a regex that only looks for one might miss the match entirely. - Ignoring whitespace/newlines: HTML often has unexpected spaces or line breaks, which can break strict regex patterns that don't account for them.
- Greedy matching: If you used a greedy
.*instead of non-greedy.*?, it might capture more than just the text inside your target tag, or fail to match correctly.
Fix 1: Adjust Your Regex
Here's a corrected regex pattern that handles the multiple classes, uses non-greedy matching, and includes the re.S flag to handle any line breaks in the HTML:
import re # Your HTML snippet html_content = '''<div class="result__links"> <p class="result__outcome u-hide-phablet">Kolkata Knight Riders won by 7 wickets</p> <p class="result__info u-hide-phablet"> Match 15, 20:00 IST (14:30 GMT), Sawai Mansingh Stadium, Jaipur </p> <a class="result__button result__button--mc btn" href="/match/2018/15?tab=sco...''' # Pattern to match the exact p tag and capture its text pattern = r'<p class="result__outcome u-hide-phablet">(.*?)</p>' match_result = re.search(pattern, html_content, re.S) if match_result: extracted_text = match_result.group(1).strip() print(extracted_text) # Output: Kolkata Knight Riders won by 7 wickets else: print("No match found")
Fix 2: Use a Proper HTML Parser (Better Long-Term Solution)
Regex is fragile for HTML parsing—even small changes to the HTML (like reordering classes or adding extra attributes) will break your pattern. A far more reliable approach is to use a library like BeautifulSoup, which is designed to parse HTML properly:
from bs4 import BeautifulSoup html_content = '''<div class="result__links"> <p class="result__outcome u-hide-phablet">Kolkata Knight Riders won by 7 wickets</p> <p class="result__info u-hide-phablet"> Match 15, 20:00 IST (14:30 GMT), Sawai Mansingh Stadium, Jaipur </p> <a class="result__button result__button--mc btn" href="/match/2018/15?tab=sco...''' # Parse the HTML soup = BeautifulSoup(html_content, 'html.parser') # Find the p tag with both target classes target_tag = soup.find('p', class_=['result__outcome', 'u-hide-phablet']) if target_tag: extracted_text = target_tag.get_text(strip=True) print(extracted_text) # Output: Kolkata Knight Riders won by 7 wickets else: print("Target tag not found")
This method is way more robust because BeautifulSoup understands HTML structure, so it won't break if the class order changes, extra whitespace is added, or other minor modifications are made to the HTML.
内容的提问来源于stack exchange,提问作者Rajat Gupta

