如何用C#提取IE渲染后htm(关联xml/xsl)的纯文本并保存为TXT
Hey Dave, let's work through how to get that plain text you need without manually copying from IE. Your HTML uses JavaScript to transform hello.xml with hello.xsl and render the result—we can replicate that logic directly in C# (either by handling the XML/XSLT ourselves, or simulating the browser load if you need full JS execution). Here are two solid approaches:
Approach 1: Direct XML + XSLT Transformation (Most Reliable)
This method skips the browser entirely and does exactly what your JavaScript does: transforms the XML with the XSLT, then extracts plain text from the resulting HTML. It's faster and doesn't depend on browser environments.
Steps & Code:
- First, install the HtmlAgilityPack NuGet package (it makes parsing HTML for plain text simple).
- Use .NET's built-in XML/XSLT classes to run the transformation, then extract the text:
using System; using System.IO; using System.Xml; using System.Xml.Xsl; using HtmlAgilityPack; class Program { static void Main() { // Update these paths if your files aren't in the same directory as the program string xmlFilePath = "hello.xml"; string xslFilePath = "hello.xsl"; string outputTxtPath = "extracted_text.txt"; // Step 1: Perform the XSLT transformation to get HTML content var xsltTransform = new XslCompiledTransform(); xsltTransform.Load(xslFilePath); string transformedHtml; using (var xmlReader = XmlReader.Create(xmlFilePath)) using (var stringWriter = new StringWriter()) { xsltTransform.Transform(xmlReader, null, stringWriter); transformedHtml = stringWriter.ToString(); } // Step 2: Parse the HTML to extract plain text var htmlDocument = new HtmlDocument(); htmlDocument.LoadHtml(transformedHtml); string plainText = htmlDocument.DocumentNode.InnerText; // Step 3: Save the plain text to a TXT file File.WriteAllText(outputTxtPath, plainText); Console.WriteLine($"Done! Plain text saved to {outputTxtPath}"); } }
Approach 2: Simulate Browser Load with WebBrowser Control
If you need to fully replicate the IE behavior (for example, if your JavaScript has extra logic beyond the XSLT transform), you can use the WinForms WebBrowser control to load the HTML, execute the JS, and extract the text.
Steps & Code:
- Create a new WinForms project (or add WinForms references to your console app).
- Use this code to load the HTML and extract text once it's fully rendered:
using System; using System.IO; using System.Windows.Forms; class HtmlTextExtractor : Form { private readonly WebBrowser _webBrowser; public HtmlTextExtractor() { _webBrowser = new WebBrowser(); _webBrowser.DocumentCompleted += OnDocumentCompleted; // Load your local HTML file string htmlPath = Path.Combine(Environment.CurrentDirectory, "hello.htm"); _webBrowser.Navigate(htmlPath); } private void OnDocumentCompleted(object sender, WebBrowserDocumentCompletedEventArgs e) { // Make sure the main document is fully loaded if (e.Url == _webBrowser.Url) { string plainText = _webBrowser.Document.Body.InnerText; File.WriteAllText("extracted_text.txt", plainText); Console.WriteLine("Done! Plain text saved to extracted_text.txt"); Close(); } } static void Main() { Application.Run(new HtmlTextExtractor()); } }
Key Notes:
- For Approach 1, HtmlAgilityPack is essential for cleanly stripping HTML tags to get plain text. Install it via NuGet Package Manager.
- Approach 2 relies on the IE rendering engine (even on newer Windows versions, the WebBrowser control uses IE in compatibility mode), so ensure IE components are available on the system where you run this.
- Always double-check that your
hello.xml,hello.xsl, andhello.htmfiles are in the correct directory (or use full file paths in the code).
内容的提问来源于stack exchange,提问作者DaveG

