You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Html Agility Pack在C#中忽略script节点提取代理IP

解决Html Agility Pack提取代理IP时的脚本节点干扰问题

我明白你现在的困扰:想从代理列表页面提取数据,结果带proxy类的li因为包含script子节点,拿不到正确的IP和端口,连País、Ciudad列的数据也获取不到。咱们一步步来解决这个问题:

问题根源分析

  1. Proxy字段为空:你之前的代码里,克隆li后直接移除了所有子节点,这会把包括脚本生成逻辑在内的所有内容都删掉,自然拿不到数据。实际上这个li里的IP和端口是通过script动态输出的,得解析脚本内容才能拿到有效信息。
  2. País/Ciudad无法获取:这些列的节点通常嵌套了图标类的子元素(比如span),直接取InnerText会混入多余内容,或者你的节点选择逻辑没有精准定位到有效文本区域。

修正后的代码方案

我调整了你的代码,主要做了这几个关键改进:

  • 用正则表达式解析script里的IP和端口内容
  • 优化了各列的文本提取逻辑,确保能拿到País、Ciudad等干净数据
using System;
using System.Collections.Generic;
using System.Linq;
using System.Text.RegularExpressions;
using HtmlAgilityPack;

class ProxyExtractor
{
    static void Main()
    {
        // 正则匹配页面脚本里的IP+端口格式
        Regex ipPortRegex = new Regex(@"document\.write\('([\d\.]+)' \+ ':' \+ '(\d+)'\)");
        
        WebClient webClient = new WebClient();
        // 若页面编码异常,可添加webClient.Encoding = System.Text.Encoding.UTF8;
        string page = webClient.DownloadString("http://proxy-list.org/spanish/search.php?search=&country=any&type=any&port=any&ssl=any");
        
        HtmlAgilityPack.HtmlDocument doc = new HtmlAgilityPack.HtmlDocument();
        doc.LoadHtml(page);
        
        List<List<string>> proxyData = doc.DocumentNode.SelectSingleNode("//div[@class='table']")
            .Descendants("ul")
            .Where(ul => ul.Elements("li").Count() >= 5) // 过滤掉不完整的代理行
            .Select(ul => 
            {
                var liElements = ul.Elements("li").ToList();
                List<string> row = new List<string>();
                
                foreach (var li in liElements)
                {
                    string result = string.Empty;
                    if (li.HasClass("proxy"))
                    {
                        // 定位script节点并解析内容
                        var scriptNode = li.SelectSingleNode("script");
                        if (scriptNode != null)
                        {
                            var scriptContent = scriptNode.InnerText.Trim();
                            var match = ipPortRegex.Match(scriptContent);
                            if (match.Success)
                            {
                                result = $"{match.Groups[1].Value}:{match.Groups[2].Value}";
                            }
                        }
                    }
                    else
                    {
                        // 清理其他列的多余空白字符,保留有效文本
                        result = string.Join(" ", li.InnerText.Split(new[] { '\n', '\r', '\t' }, StringSplitOptions.RemoveEmptyEntries)).Trim();
                    }
                    row.Add(result);
                }
                return row;
            }).ToList();
        
        // 测试输出提取结果
        foreach (var row in proxyData)
        {
            Console.WriteLine($"Proxy: {row[0]}, País: {row[1]}, Ciudad: {row[2]}, Tipo: {row[3]}, Velocidad: {row[4]}, HTTPS/SSL: {row[5]}");
        }
    }
}

关键改进点说明

  1. 正则解析脚本:针对页面里document.write('IP' + ':' + 'Port')的固定脚本格式,用正则精准提取IP和端口,彻底避开脚本内容的干扰。
  2. 文本清理逻辑:对País、Ciudad等列的文本做了空白字符过滤,确保拿到的是干净的有效数据。
  3. 行有效性过滤:通过判断li元素的数量,只处理包含完整代理信息的行,避免无效节点混入结果。

如果后续页面的脚本格式有变化,你只需要调整正则表达式的匹配规则就能快速适配。

内容的提问来源于stack exchange,提问作者Willy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:07:58