使用ScrapeGraphAI爬取Appian论坛仅获首条评论,求获取全部回复方案
问题
我有一个Appian社区论坛管理板块的帖子URL列表,用ScrapeGraphAI爬取时只能获取每个帖子的第一条用户评论,拿不到后续回复;之前用BS4尝试也失败了。以下是我当前的Python代码:
import json from scrapegraphai.graphs import SmartScraperGraph def main(): graph_config = { "llm": { "model": "ollama/llama3", "temperature": 0, "base_url": "http://localhost:11434", "format": "json", # Ollama needs the format to be specified explicitly }, "embeddings": { "model": "ollama/nomic-embed-text", } } source_urls = [] with open('cleaned-urls.txt', 'r') as f: sources = [line.strip() for line in f if line.strip()] source_urls.extend(sources) for source_url in source_urls: try: prompt = "find the best way to extract data, eliminate unneccesary fields and organise to only show the entire conversation and code snippets. make sure to include all text from the conversation and the users answers. always the first text is the question and what follows is from other user replies" smart_scraper_graph = SmartScraperGraph(prompt=prompt, source=source_url, config=graph_config) result = smart_scraper_graph.run() output = json.dumps(result, indent=2) print(output) except Exception as e: print(f"An error occurred: {e}") if __name__ == "__main__": main()
URL列表:
- https://community.appian.com/discussions/f/administration/14/integrate-token-device-with-appian
- https://community.appian.com/discussions/f/administration/27/how-do-we-configure-enable-appian-tempo
- https://community.appian.com/discussions/f/administration/31/how-to-download-get-pdf-of-the-documentation
- https://community.appian.com/discussions/f/administration/39/we-need-to-establish-a-single-signon-with-an-outside-web-site-for-which-a-certif
- https://community.appian.com/discussions/f/administration/43/is-there-a-way-to-import-an-application-exported-from-appian-6-6-1and-import-in
- https://community.appian.com/discussions/f/administration/47/we-are-having-issues-with-oracle-db-integration-for-appian-6-6-1-we-are-install
解决方法
1. 优化ScrapeGraphAI的Prompt与配置
- 细化Prompt指令:原Prompt不够明确,需指定评论的层级和范围,修改为:
提取页面中所有对话内容:包括楼主的提问、所有楼层的用户回复(含嵌套回复)及所有代码片段。按对话顺序整理为结构化数据:{"提问": "...", "回复": [{"用户": "...", "内容": "...", "代码": "..."}...]},确保不遗漏任何一条回复。 - 启用动态渲染:Appian论坛的评论可能是动态加载的,需在配置中添加渲染选项:
修改graph_config,新增scraper配置:graph_config = { "llm": { "model": "ollama/llama3", "temperature": 0, "base_url": "http://localhost:11434", "format": "json", }, "embeddings": { "model": "ollama/nomic-embed-text", }, "scraper": { "use_playwright": True, # 启用Playwright渲染动态内容 "headless": False, # 调试时可设为False查看浏览器窗口 } }
2. 改用BS4配合Playwright处理动态内容
如果ScrapeGraphAI仍有局限,可直接用Playwright加载页面后再用BS4解析:
from bs4 import BeautifulSoup from playwright.sync_api import sync_playwright import json def scrape_appian_thread(url): with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto(url) # 等待评论区加载完成,可根据页面元素调整等待条件 page.wait_for_selector('.discussion-post') # 滚动页面加载所有懒加载评论 for _ in range(3): page.evaluate('window.scrollTo(0, document.body.scrollHeight)') page.wait_for_timeout(1000) html = page.content() browser.close() soup = BeautifulSoup(html, 'html.parser') posts = soup.find_all('div', class_='discussion-post') conversation = { "提问": posts[0].find('div', class_='post-body').get_text(strip=True), "回复": [] } for post in posts[1:]: user = post.find('span', class_='user-name').get_text(strip=True) content = post.find('div', class_='post-body').get_text(strip=True) code_snippets = [code.get_text(strip=True) for code in post.find_all('code')] conversation["回复"].append({ "用户": user, "内容": content, "代码片段": code_snippets if code_snippets else None }) return conversation def main(): source_urls = [] with open('cleaned-urls.txt', 'r') as f: source_urls = [line.strip() for line in f if line.strip()] for url in source_urls: try: result = scrape_appian_thread(url) print(json.dumps(result, indent=2, ensure_ascii=False)) except Exception as e: print(f"处理URL {url}时出错: {e}") if __name__ == "__main__": main()
需先安装依赖:pip install beautifulsoup4 playwright,再运行playwright install安装浏览器驱动。
3. 直接调用论坛API接口
打开浏览器开发者工具(F12),查看网络请求,找到加载评论的API接口(通常是XHR/Fetch请求),直接调用该接口获取结构化的评论数据,这种方式效率更高、稳定性更强。
内容的提问来源于stack exchange,提问作者Adrian Puiu
相关产品推荐
相关产品推荐

