You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Swift技术问题:For循环中WKWebView加载多站HTML为空如何解决?

Fixing Empty HTML Issues with WKWebView Course Scraping

Hey there! Let's break down why you're getting empty HTML documents when trying to scrape those UMD course pages, and fix it step by step.

The Core Problems in Your Current Code

  • WKWebView loads content asynchronously: You can't just create a URL and expect immediate access to the HTML—you need to wait for the page to fully load first.
  • No actual web view loading in createTapped: Right now, you're just creating URL objects but never telling the web view to load them.
  • Missing loadURL implementation: You call this function in your add action, but it's not defined anywhere.
  • No delegate to track load completion: Without a navigation delegate, you have no way to know when the page is ready to be parsed.

Step-by-Step Fix

First, let's update your ViewController to handle asynchronous loading properly:

import UIKit
import WebKit

class ViewController: UIViewController, WKNavigationDelegate { // Add WKNavigationDelegate conformance
    @IBOutlet weak var classCode: UITextField!
    var webview = WKWebView()
    var classCodes: [String] = []
    var currentCodeIndex = 0 // Track which course code we're processing

    override func viewDidLoad() {
        super.viewDidLoad()
        // Set up the web view's delegate and add it to the view (hidden if needed)
        webview.navigationDelegate = self
        webview.frame = view.bounds
        webview.isHidden = true // Hide it since we don't need to display the page
        view.addSubview(webview)
    }

    @IBAction func add() {
        guard let code = classCode.text, !code.isEmpty else {
            // Add validation here if needed (check for valid course code format)
            return
        }
        classCodes.append(code)
        classCode.text = nil
    }

    @IBAction func createTapped() {
        guard !classCodes.isEmpty else { return }
        currentCodeIndex = 0
        // Start loading the first course page
        loadCoursePage(for: classCodes[currentCodeIndex])
    }

    private func loadCoursePage(for code: String) {
        guard let url = URL(string: "https://app.testudo.umd.edu/soc/search?courseId=\(code)&sectionId=&termId=201801&openSectionsOnly=true&_openSectionsOnly=on&creditCompare=&credits=&courseLevelFilter=ALL&instructor=&_facetoface=on&_blended=on&_online=on&courseStartCompare=&courseStartHour=&courseStartMin=&courseStartAM=&courseEndHour=&courseEndMin=&courseEndAM=&teachingCenter=") else {
            // Handle invalid URL (maybe code has invalid characters?)
            print("Invalid URL for course code: \(code)")
            proceedToNextCourse()
            return
        }
        webview.load(URLRequest(url: url))
    }

    // MARK: - WKNavigationDelegate
    func webView(_ webView: WKWebView, didFinish navigation: WKNavigation!) {
        // Now the page is loaded—fetch the HTML
        webView.evaluateJavaScript("document.documentElement.outerHTML") { [weak self] result, error in
            guard let self = self else { return }
            if let html = result as? String {
                print("HTML for \(self.classCodes[self.currentCodeIndex]):\n\(html)")
                // Here's where you'd parse the HTML (use a parser like SwiftSoup if needed)
                self.parseHTML(html: html, for: self.classCodes[self.currentCodeIndex])
            } else if let error = error {
                print("Error fetching HTML: \(error.localizedDescription)")
            }
            // Move to the next course code
            self.proceedToNextCourse()
        }
    }

    private func proceedToNextCourse() {
        currentCodeIndex += 1
        if currentCodeIndex < classCodes.count {
            loadCoursePage(for: classCodes[currentCodeIndex])
        } else {
            print("All course pages processed!")
        }
    }

    private func parseHTML(html: String, for code: String) {
        // Add your HTML parsing logic here
        // Example using SwiftSoup (you'd need to add it via CocoaPods/Swift Package Manager):
        // do {
        //     let doc = try SwiftSoup.parse(html)
        //     let sections = try doc.select(".section").array()
        //     // Process sections...
        // } catch {
        //     print("Parsing error: \(error)")
        // }
    }
}

Key Changes Explained

  1. Added WKNavigationDelegate conformance: This lets us listen for when the web view finishes loading a page.
  2. Web view setup: We added the web view to the view hierarchy (hidden, since we don't need to display it) and set its delegate.
  3. Sequential loading: We track the current course index and load one page at a time, moving to the next only after the current one is processed.
  4. evaluateJavaScript to get HTML: WKWebView doesn't expose HTML directly, so we use JavaScript to fetch the full document HTML once the page loads.
  5. Validation checks: Added guards to prevent empty codes and invalid URLs.

Additional Tips

  • HTML Parsing: For easier parsing, consider using a library like SwiftSoup instead of manually parsing the HTML string.
  • Rate Limiting: Be respectful to the website—add small delays between requests if needed to avoid getting blocked.
  • Error Handling: Expand the error handling to catch cases where the page fails to load or returns a non-200 status code.

内容的提问来源于stack exchange,提问作者Edward_Snowden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:49:01