如何用Perl从Ask.com特定标签中精准提取域名URL
The issue with your current code is that it processes every <p> tag in the HTML, which includes unrelated content like "Videos" or "Change Settings". To fix this, we need to filter <p> tags to only those with the class PartialSearchResults-item-url, and optionally validate that the content is a valid domain URL.
Here's the corrected code:
use HTML::TokeParser; my $p = HTML::TokeParser->new(\$rrs); # Open output file once (more efficient than opening/closing in loop) open(my $out_fh, '>>', 'askurls.txt') or die "Could not open file: $!"; while (my $tag_info = $p->get_tag("p")) { # Extract attributes from the <p> tag my %tag_attrs = %{$tag_info->[1]}; # Check if the tag has the target class if (exists $tag_attrs{class} && $tag_attrs{class} eq 'PartialSearchResults-item-url') { # Get trimmed text inside the <p> tag my $url = $p->get_trimmed_text("/p"); # Optional: Validate the content is a valid domain (avoids any stray text) if ($url =~ /^[a-z0-9.-]+\.[a-z]{2,}$/i) { print "$url\n"; print $out_fh "$url\n"; } } } close($out_fh);
Key Changes Explained:
- Filter by Tag Class: We now check if the
<p>tag has the exact classPartialSearchResults-item-urlbefore extracting its text. This eliminates all unrelated<p>tags. - Efficient File Handling: Instead of opening and closing the output file inside the loop (which is slow for large datasets), we open it once before processing and close it afterward.
- Optional URL Validation: The regex
/^[a-z0-9.-]+\.[a-z]{2,}$/iensures we only keep valid domain URLs (likemaps.google.com) and skip any unexpected text that might slip through. - Scalar Variable for URL:
get_trimmed_text("/p")returns a single string, so using a scalar ($url) is cleaner than an array.
This code will now only extract the domain URLs from the specific <p> tags you're targeting, ignoring all other content.
内容的提问来源于stack exchange,提问作者Haroon Ahmad
相关产品推荐
相关产品推荐

