You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用C# HttpWebRequest获取表格数据及XPath定位指定TD文本问题

解决C#中获取指定表格数据的问题

咱们一步步来搞定这个需求:先用HttpWebRequest拿到网页内容,再精准定位到你要的那个td元素的文本。直接上实用方案:

第一步:用HttpWebRequest获取网页HTML内容

先写个工具方法获取目标链接的完整HTML,记得加个UserAgent避免被网站拦截:

using System;
using System.IO;
using System.Net;

public class HtmlFetcher
{
    public static string GetHtml(string targetUrl)
    {
        HttpWebRequest request = (HttpWebRequest)WebRequest.Create(targetUrl);
        request.Method = "GET";
        // 模拟浏览器请求头,降低被拦截概率
        request.UserAgent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36";
        
        using (HttpWebResponse response = (HttpWebResponse)request.GetResponse())
        using (Stream responseStream = response.GetResponseStream())
        using (StreamReader reader = new StreamReader(responseStream))
        {
            return reader.ReadToEnd();
        }
    }
}

第二步:解析HTML提取目标td内容

原生的XmlDocument处理HTML容易踩坑(毕竟很多网页的HTML不规范),我强烈推荐用HtmlAgilityPack这个NuGet包来解析,它对各种不规范HTML的兼容性拉满。

1. 先安装HtmlAgilityPack

打开NuGet包管理器搜索HtmlAgilityPack安装,或者用命令行:

Install-Package HtmlAgilityPack

2. 用XPath精准定位目标元素

你已经会定位包含指定文本的a标签,现在只需要调整XPath路径,找到它后面的第二个td。对应的XPath表达式是:

//a[contains(., 'Example String')]/../following-sibling::td[2]

给你拆解下这个路径:

  • //a[contains(., 'Example String')]:定位到包含目标文本的a标签
  • /../:跳转到a标签的父元素(也就是包裹它的那个td)
  • following-sibling::td[2]:找到这个父td后面的第二个同级td元素(正好是你要的89所在的td)

3. 完整解析代码

把获取到的HTML传入,就能拿到目标文本:

using HtmlAgilityPack;

public class HtmlParser
{
    public static string GetTargetTdContent(string html)
    {
        HtmlDocument doc = new HtmlDocument();
        doc.LoadHtml(html);
        
        // 替换成你实际要匹配的字符串
        string targetString = "Example String";
        string xpath = $"//a[contains(., '{targetString}')]/../following-sibling::td[2]";
        
        HtmlNode targetTd = doc.DocumentNode.SelectSingleNode(xpath);
        
        // 处理找不到元素的情况,避免空引用
        return targetTd?.InnerText.Trim() ?? "未找到目标内容";
    }
}

组合起来用

把两部分代码结合,就能完成整个流程:

string targetUrl = "你的目标网页链接";
string pageHtml = HtmlFetcher.GetHtml(targetUrl);
string result = HtmlParser.GetTargetTdContent(pageHtml);
Console.WriteLine(result); // 输出89

如果网站有反爬机制,可能还需要处理Cookie、代理或者请求频率限制,但基础的核心流程就是这样啦!

内容的提问来源于stack exchange,提问作者Uni VPS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:52:46