如何用Java统计HTTP响应体中各HTML标签出现次数并按次数排序
Alright, let's break this down into actionable steps with Java code. I'll walk you through how to count HTML tag occurrences in your HTTP response string and sort them by frequency.
Solution Overview
First, we need to tackle three core tasks:
- Extract all HTML tags from the input string (handling opening, closing, and self-closing tags)
- Count how many times each tag appears (treating
<div>and</div>as the same tag) - Sort the tags by their occurrence count (we'll use descending order by default)
Java Implementation
Here's a complete, runnable example tailored to your needs:
import java.util.*; import java.util.regex.Matcher; import java.util.regex.Pattern; public class HtmlTagCounter { public static void main(String[] args) { // Your input HTTP response body string String htmlContent = "<div><p><span class=\"lede\">Today, the European </span>Space Agency&apos;s Rosetta spacecraft will engage its thrusters for one final maneuver: a suicidal plunge toward the comet it has been orbiting for two years and chasing for a decade. After Rosetta collides with comet 67P/Churyumov-Gerasimenko..."; // Step 1: Regex to match HTML tags (captures tag name, ignores attributes and closing slashes) Pattern tagPattern = Pattern.compile("<(/?)(\\w+)[^>]*>"); Matcher matcher = tagPattern.matcher(htmlContent); // Step 2: Count tag occurrences Map<String, Integer> tagCountMap = new HashMap<>(); while (matcher.find()) { // Normalize tag name to lowercase to handle case variations like <DIV> vs <div> String tagName = matcher.group(2).toLowerCase(); // Update count: increment if tag exists, set to 1 if it's new tagCountMap.put(tagName, tagCountMap.getOrDefault(tagName, 0) + 1); } // Step 3: Sort tags by occurrence count (descending order) List<Map.Entry<String, Integer>> sortedTags = new ArrayList<>(tagCountMap.entrySet()); sortedTags.sort((entry1, entry2) -> entry2.getValue().compareTo(entry1.getValue())); // Step 4: Print the results System.out.println("HTML Tag Occurrences (sorted by count):"); for (Map.Entry<String, Integer> entry : sortedTags) { System.out.printf("%s: %d%n", entry.getKey(), entry.getValue()); } } }
Key Details Explained
Let's unpack what each part does:
- Regex Matching: The pattern
<(/?)(\\w+)[^>]*>is designed to:- Capture the optional closing slash (
/?) to identify closing tags - Grab the tag name (
\\w+) which includes letters, numbers, and underscores - Ignore any attributes (like
class="lede") with[^>]*(all characters until the closing>)
- Capture the optional closing slash (
- Counting: Using a
HashMaplets us efficiently track tag frequencies. Converting tag names to lowercase ensures consistency even if tags are written in mixed case. - Sorting: We convert the map entries to a list, then use a comparator to sort by count in descending order. Swap
entry2.getValue()andentry1.getValue()if you want ascending order instead. - Edge Cases: This code handles opening tags, closing tags, self-closing tags (like
<br/>), and case variations seamlessly.
Sample Output
For your specific input string, the output will look like this:
HTML Tag Occurrences (sorted by count): div: 1 p: 1 span: 1
If your input had repeated tags (e.g., multiple <span> elements or closing tags like </div>), the counts would update automatically.
内容的提问来源于stack exchange,提问作者Tomek Bieniek
相关产品推荐
相关产品推荐

