You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在BeautifulSoup中用正则提取<p>标签文本?Python脚本返回空列表

Fixing Your Empty List Issue When Extracting Text from a <p> Tag with Regex

Hey Rajat, let's break down why your Python regex is returning an empty list and get you the text you need from that <p> tag.

Common Regex Pitfalls Here

The most likely issues with your original regex are:

  • Not accounting for multiple class values: Your target <p> has two classes (result__outcome and u-hide-phablet), so a regex that only looks for one might miss the match entirely.
  • Ignoring whitespace/newlines: HTML often has unexpected spaces or line breaks, which can break strict regex patterns that don't account for them.
  • Greedy matching: If you used a greedy .* instead of non-greedy .*?, it might capture more than just the text inside your target tag, or fail to match correctly.

Fix 1: Adjust Your Regex

Here's a corrected regex pattern that handles the multiple classes, uses non-greedy matching, and includes the re.S flag to handle any line breaks in the HTML:

import re

# Your HTML snippet
html_content = '''<div class="result__links"> <p class="result__outcome u-hide-phablet">Kolkata Knight Riders won by 7 wickets</p> <p class="result__info u-hide-phablet"> Match 15, 20:00 IST (14:30 GMT), Sawai Mansingh Stadium, Jaipur </p> <a class="result__button result__button--mc btn" href="/match/2018/15?tab=sco...'''

# Pattern to match the exact p tag and capture its text
pattern = r'<p class="result__outcome u-hide-phablet">(.*?)</p>'
match_result = re.search(pattern, html_content, re.S)

if match_result:
    extracted_text = match_result.group(1).strip()
    print(extracted_text)  # Output: Kolkata Knight Riders won by 7 wickets
else:
    print("No match found")

Fix 2: Use a Proper HTML Parser (Better Long-Term Solution)

Regex is fragile for HTML parsing—even small changes to the HTML (like reordering classes or adding extra attributes) will break your pattern. A far more reliable approach is to use a library like BeautifulSoup, which is designed to parse HTML properly:

from bs4 import BeautifulSoup

html_content = '''<div class="result__links"> <p class="result__outcome u-hide-phablet">Kolkata Knight Riders won by 7 wickets</p> <p class="result__info u-hide-phablet"> Match 15, 20:00 IST (14:30 GMT), Sawai Mansingh Stadium, Jaipur </p> <a class="result__button result__button--mc btn" href="/match/2018/15?tab=sco...'''

# Parse the HTML
soup = BeautifulSoup(html_content, 'html.parser')

# Find the p tag with both target classes
target_tag = soup.find('p', class_=['result__outcome', 'u-hide-phablet'])

if target_tag:
    extracted_text = target_tag.get_text(strip=True)
    print(extracted_text)  # Output: Kolkata Knight Riders won by 7 wickets
else:
    print("Target tag not found")

This method is way more robust because BeautifulSoup understands HTML structure, so it won't break if the class order changes, extra whitespace is added, or other minor modifications are made to the HTML.

内容的提问来源于stack exchange,提问作者Rajat Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:02:07