在Java字符串中移除HTML表格最后列为N/A的行
Hey there! Let's solve this problem of stripping out all table rows where the last column contains "N/A" from your HTML string (stored as a Java String).
Regex Pattern Explanation
We need a regex that matches entire <tr> elements where the final data column (a <td> tag) contains exactly N/A—accounting for possible whitespace, newlines, or minor formatting variations common in HTML.
Here's the Java-friendly pattern (note the double backslashes required for escaping in Java strings):
String regex = "<tr[^>]*>(?:\\s*<td[^>]*>.*?</td>\\s*)+\\s*<td[^>]*>\\s*N/A\\s*</td>\\s*</tr>";
Breakdown of the pattern:
<tr[^>]*>: Matches the opening<tr>tag, allowing any attributes likeclassorid.(?:\\s*<td[^>]*>.*?</td>\\s*)+: Non-capturing group that matches one or more preceding<td>columns (with any content inside), accounting for surrounding whitespace/newlines.\\s*<td[^>]*>\\s*N/A\\s*</td>\\s*: Matches the final<td>that containsN/A(with optional whitespace before/after the text to handle messy formatting).</tr>: Matches the closing</tr>tag to capture the entire row.
Java Code Example
You can use String.replaceAll() for quick implementation, or Pattern/Matcher for more control (like case-insensitive matching):
Basic Implementation with replaceAll()
// Your original HTML string (truncated example expanded for clarity) String html = "<table class=\"overviewTable\"> <tr> <th colspan=\"6\" class=\"header suite\"> <div class=\"suiteLinks\"> <a href=\"suite1_groups.html\">Groups</a> </div> Test Automation </th> </tr> <tr class=\"columnHeadings\"> <td> </td> <th>Duration</th> <th>Status</th> <th>Tests</th> <th>Errors</th> <th>Failures</th> </tr> <tr> <td>Test 1</td> <td>10s</td> <td>Passed</td> <td>5</td> <td>0</td> <td>N/A</td> </tr> <tr> <td>Test 2</td> <td>15s</td> <td>Passed</td> <td>7</td> <td>0</td> <td>2</td> </tr>"; // Remove rows with N/A in the last column String cleanedHtml = html.replaceAll(regex, ""); System.out.println(cleanedHtml);
Advanced Implementation with Pattern/Matcher
If you need to handle case variations (like n/a or N/a), add the CASE_INSENSITIVE flag. The DOTALL flag ensures the regex matches across line breaks:
Pattern pattern = Pattern.compile(regex, Pattern.CASE_INSENSITIVE | Pattern.DOTALL); Matcher matcher = pattern.matcher(html); String cleanedHtml = matcher.replaceAll("");
Important Notes
- This regex assumes data rows use
<td>tags for columns—header rows with<th>won't be affected, which is exactly what we want. - If your
<td>tags contain nested elements (like links or spans), the pattern will still work as long as the final column's visible text isN/A. - For extremely complex HTML structures, a dedicated parser like Jsoup is more reliable. But for the specific table structure you provided, this regex will work perfectly.
内容的提问来源于stack exchange,提问作者Mukesh Rajput

