You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式替换格式异常的</td>标签

Fixing Malformed Tags with Python Regex

Got it, let's work through this problem. You need to hunt down those messed-up variants where the t and d are out of order, there's a random space, and maybe an equals sign thrown in—like </td= >, </t= d>, or </= td>—and replace them all with the standard </td>.

Here's a straightforward solution using Python's re module:

Step 1: The Regex Pattern

First, let's build a regex that targets exactly these malformed tags. The key is to ensure the tag contains:

  • The opening </
  • Both t and d characters (regardless of order)
  • At least one space
  • Any extra characters like = (since your examples include them)
  • The closing >

The pattern we'll use is:

r'</(?=[^>]*t)(?=[^>]*d)(?=[^>]* )[^>]*>'

Breakdown of the Pattern:

  • </: Matches the start of the closing tag
  • (?=[^>]*t): Positive lookahead to confirm the tag contains a t somewhere before the >
  • (?=[^>]*d): Another positive lookahead to confirm the tag contains a d
  • (?=[^>]* ): Lookahead to ensure there's at least one space in the tag
  • [^>]*: Matches any characters except > (covers all the messy stuff in between)
  • >: Matches the end of the tag

Step 2: Python Code Implementation

Here's a reusable function to clean up your text:

import re

def repair_td_tags(input_text):
    # Regex to match malformed </td> variants
    malformed_td_pattern = r'</(?=[^>]*t)(?=[^>]*d)(?=[^>]* )[^>]*>'
    # Replace all matches with standard </td>
    return re.sub(malformed_td_pattern, '</td>', input_text)

# Test it out with your examples
sample_text = """
<table>
    <tr>
        <td>Valid content</td= >
        <td>More content</t= d>
        <td>Even more</= td>
    </tr>
</table>
"""

cleaned_text = repair_td_tags(sample_text)
print(cleaned_text)

Step 3: Output of the Test

Running the code above will give you this cleaned output:

<table>
    <tr>
        <td>Valid content</td>
        <td>More content</td>
        <td>Even more</td>
    </tr>
</table>

Optional: Restrict to Only Allowed Characters

If you want to make sure we don't accidentally match other tags that happen to have t, d, and a space (unlikely, but safe), you can modify the regex to only allow t, d, =, and spaces in the tag:

r'</(?=[td= ]*t)(?=[td= ]*d)(?=[td= ]* )[td= ]*>'

This version will ignore any tags with extra characters (like </tdx >) that don't fit your specific malformed pattern.

内容的提问来源于stack exchange,提问作者krishna bharadia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:02:35