You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Unity中C#脚本提取网页HTML表格数据的实现方法咨询

提取HTML中指定表格数据的Unity C#实现思路

嘿,针对你在Unity里要提取网页指定表格数据的需求,我整理了几个从简单到健壮的实现思路,你可以根据自己的场景挑合适的:

1. 简单字符串截取法

如果目标网页的HTML结构非常固定,这个方法最直接高效,直接定位目标表格的起始标记后截取内容:

// 假设你已经获取到完整的网页源码字符串htmlContent
string startMarker = "<table class=\"formular\">";
int startIndex = htmlContent.IndexOf(startMarker);

if (startIndex != -1)
{
    // 先截取从起始标记到末尾的内容
    string tableContent = htmlContent.Substring(startIndex);
    
    // 可选:如果只需要表格本身,截取到</table>结束标记为止
    int endIndex = tableContent.IndexOf("</table>") + "</table>".Length;
    if (endIndex > 0)
    {
        tableContent = tableContent.Substring(0, endIndex);
    }
    
    // 保存提取后的内容到文件
    File.WriteAllText(Application.persistentDataPath + "/extracted_table.html", tableContent);
}
else
{
    Debug.LogError("未找到目标表格的起始标记");
}

优点:零依赖,代码量极少,适合结构完全固定的页面;缺点:对HTML格式变化极度敏感,比如class名多了空格、大小写变化就会失效。

2. 正则表达式匹配法

如果字符串截取的灵活性不够,可以用正则来匹配表格区域,能容忍一些格式上的小差异:

using System.Text.RegularExpressions;

// 正则匹配class为formular的table标签及其内部所有内容
string pattern = @"<table\s+class=""formular""[^>]*>.*?</table>";
Match tableMatch = Regex.Match(htmlContent, pattern, RegexOptions.Singleline);

if (tableMatch.Success)
{
    string tableHtml = tableMatch.Value;
    File.WriteAllText(Application.persistentDataPath + "/extracted_table.html", tableHtml);
}
else
{
    Debug.LogError("正则表达式未匹配到目标表格");
}

说明:RegexOptions.Singleline参数让正则的.能匹配换行符,确保跨行匹配整个表格。
优点:比字符串截取更灵活,能处理标记内的额外空格;缺点:HTML结构复杂时(比如嵌套表格)正则容易匹配出错,毕竟正则不是专门为解析HTML设计的工具。

3. 专业HTML解析库(推荐)

如果需要处理结构复杂或不规范的HTML,强烈推荐使用HtmlAgilityPack——这是一个成熟的.NET HTML解析库,能完美兼容Unity,处理各种奇葩的HTML代码:

操作步骤:

  1. 下载适合.NET Framework的HtmlAgilityPack DLL,导入到Unity项目的Plugins文件夹中;
  2. 使用以下代码提取表格:
using HtmlAgilityPack;

HtmlDocument htmlDoc = new HtmlDocument();
htmlDoc.LoadHtml(htmlContent);

// 用XPath查找class为formular的table节点
HtmlNode targetTable = htmlDoc.DocumentNode.SelectSingleNode("//table[@class='formular']");

if (targetTable != null)
{
    // 获取表格的完整HTML代码
    string tableHtml = targetTable.OuterHtml;
    File.WriteAllText(Application.persistentDataPath + "/extracted_table.html", tableHtml);
    
    // 进阶:如果要提取表格内的具体数据(比如转成CSV),可以遍历节点
    foreach (HtmlNode row in targetTable.SelectNodes(".//tr"))
    {
        List<string> rowData = new List<string>();
        // 遍历表头和单元格
        foreach (HtmlNode cell in row.SelectNodes(".//th|.//td"))
        {
            rowData.Add(cell.InnerText.Trim());
        }
        // 将每行数据转成CSV格式(示例)
        string csvLine = string.Join(",", rowData);
        // 这里可以写入CSV文件或做其他处理
    }
}
else
{
    Debug.LogError("未找到目标表格节点");
}

优点:完全解析HTML结构,能处理各种不规范代码,还能轻松提取表格内的具体数据;缺点:需要导入外部库,但这个库非常稳定,几乎没有兼容性问题。

内容的提问来源于stack exchange,提问作者RoyalCoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:58:03