You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于LLM的网页爬虫中无URL变化分页的处理方案咨询

基于LLM的网页爬虫中无URL变化分页的处理方案咨询

Hey there! Great question—handling dynamic pagination without distinct URLs is a super common pain point when building a generic web scraper, especially when you’re targeting a bunch of diverse sites. You’re right that hardcoding Selenium selectors for each site isn’t feasible, so here are some flexible, LLM-powered approaches that adapt to different sites automatically:

  • LLM-driven detection of pagination controls
    Use a headless browser (like Puppeteer or Playwright) to render the full page and grab its DOM structure. Then feed a trimmed-down version of that HTML to your LLM with a prompt like:

    "Look through this HTML snippet and identify all elements that handle pagination—this includes 'next page' buttons, page number links, 'load more' buttons, or infinite scroll triggers. Return their CSS selectors or XPath expressions, and note what action (click, scroll) triggers the next page."
    The LLM will pick up on common patterns (like classes named pagination-next or text like "Next →") even across totally different site designs. Just make sure to limit the HTML you send to avoid hitting context limits—focus on the main content area rather than the entire page.

  • Analyze network requests with LLM
    Most dynamic pagination loads new content via AJAX or API calls, not full page refreshes. Use your headless browser’s network interceptor to capture all requests made when the page loads (or when you interact with a pagination element). Then send the list of request URLs, parameters, and response snippets to your LLM with a prompt like:

    "Which of these requests are used to load additional pages of content? What parameters (like page, offset, limit) control the page number, and how do you increment them to get the next set of results?"
    This lets you bypass the frontend entirely—you can directly call the pagination API with the right parameters, which is faster and more reliable than simulating clicks. The LLM can even help spot required headers or session cookies that need to be included in subsequent requests.

  • Generate dynamic interaction scripts with LLM
    If you do need to simulate user interactions (like clicking a next button or scrolling for infinite load), let the LLM write the code for you. After rendering the page, send the DOM or a text description of the page to the LLM with a prompt like:

    "Write a Puppeteer script to navigate to the next page of content on this site. Use the DOM structure provided to target the correct element, and include basic error handling in case the element isn’t found."
    You can execute this generated script on the fly—no need to hardcode selectors. If the script fails (say, the site’s UI changed), just feed the error message and updated DOM back to the LLM to get a revised script.

  • Content-based pagination logic
    For sites with no obvious pagination controls (like infinite scroll or content that cuts off mid-list), use the LLM to analyze the page content itself. Send the text of the current page and the previous page to the LLM with a prompt like:

    "Compare these two page contents. Is there evidence that there’s more content to load? Look for phrases like 'Showing 1-20 of 100 results' or repeated content patterns that suggest truncation. If so, what action (scroll to bottom, click a hidden load button) would likely load the next set?"
    This works well for sites that don’t expose clear pagination UI but hint at more content through text or structure.

The core idea here is to leverage the LLM’s strength in parsing unstructured, site-specific data (HTML, network logs, content text) alongside generic browser automation tools. This way, your scraper adapts to each site’s unique pagination logic without you having to write custom code for every single target.

备注:内容来源于stack exchange,提问作者Arash Vahabpour

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 12:19:30