如何用正则表达式抓取含可选<span>的<p>标签及指定HTML表格?
Hey there! Let's break down how to handle your HTML parsing needs with regex, though I’ll start with a quick caveat: regex isn’t the best tool for complex HTML (since HTML isn’t a regular language), but for your specific use cases, these patterns should work reliably. Here’s what you need:
1. Grab <p> Tags with Optional <span> Inside
This regex will match <p> tags that may contain a <span> (with any attributes) or just plain text, and capture the core content inside the <p>:
<p>(?:<span[^>]*>)?(.*?)(?:<\/span>)?<\/p>
Pattern Breakdown:
<p>: Matches the opening<p>tag exactly(?:<span[^>]*>)?: Optional non-capturing group for the opening<span>tag (handles any span attributes like styles or classes)(.*?): Captures the inner content (uses non-greedy matching to avoid accidentally grabbing multiple tags)(?:<\/span>)?: Optional non-capturing group for the closing</span>tag<\/p>: Matches the closing</p>tag exactly
2. Parse Your Specific HTML Table Structure
For the table you shared (with mixed <p>/<span> usage in table cells), here’s a regex that captures each row’s title and text pairs, even when some <p> tags are incomplete:
<tr>\s*<td[^>]*>(?:<p>(?:<span[^>]*>)?(.*?)(?:<\/span>)?<\/p>)?<\/td>\s*<td[^>]*>(?:<p>(?:<span[^>]*>)?(.*?)(?:<\/span>)?<\/p>)?<\/td>\s*<\/tr>
Pattern Breakdown:
<tr>\s*: Matches the opening<tr>tag plus any surrounding whitespace<td[^>]*>: Matches<td>tags with any attributes (like thewidthvalues in your example)- The inner
<p>/<span>logic mirrors the first pattern, wrapped in a non-capturing group with?to handle cases where<p>tags are missing or incomplete (like your second title example) (.*?): Captures the title (from the first<td>) and text (from the second<td>) separately\s*<\/td>: Matches closing</td>tags plus any surrounding whitespace- Repeats for the second
<td>, then matches the closing</tr>tag
Example Usage (Python)
If you’re working in Python, here’s how you can extract the data using this regex:
import re # Your sample HTML html_content = """<table><tr> <td width="100"><p><span style=" font-family:'MS Shell Dlg 2'; font-size:8.25pt;">My title example:</span></p></td> <td width="440"><p><span style=" font-family:'MS Shell Dlg 2'; font-size:8.25pt;">My text example.</span></p></td> </tr> <tr> <td width="100">My second title:</p></td></tr>""" # Compile the regex pattern (re.DOTALL allows . to match newlines) table_pattern = re.compile(r'<tr>\s*<td[^>]*>(?:<p>(?:<span[^>]*>)?(.*?)(?:<\/span>)?<\/p>)?<\/td>\s*<td[^>]*>(?:<p>(?:<span[^>]*>)?(.*?)(?:<\/span>)?<\/p>)?<\/td>\s*<\/tr>', re.DOTALL) # Extract all matches row_matches = table_pattern.findall(html_content) # Print the results for title, text in row_matches: print(f"Title: {title.strip()}") if text: print(f"Text: {text.strip()}\n")
Sample Output:
Title: My title example: Text: My text example. Title: My second title: Text:
Quick Reminder
While these regex patterns work for your specific structure, they’ll struggle if the HTML gets more complex (like nested tags, inconsistent formatting, or unexpected attributes). For production code, I’d strongly recommend using an HTML parser like BeautifulSoup (Python) or jsdom (JavaScript) instead—they’re built to handle HTML’s quirks properly.
内容的提问来源于stack exchange,提问作者Saelyth

