如何用Python Playwright滚动TikTok页面至第n个帖子并抓取HTML
TikTok用户页面滚动至指定数量可见帖子的问题解决
我需要在TikTok用户页面(例如@thebeatles的页面)滚动到至少72个帖子可见,再保存HTML内容。该页面当前共有73个帖子,但我用page.keyboard.press("End")按键滚动,仅能获取到30个帖子。抓取时已中止图片和字体请求以优化性能,相关代码如下,请问如何实现滚动到第72个可见帖子?
原代码
test_fetch_raw_source.py
def test_fetch_front_page_with_n_latest_posts(): # Fetch front page with n_count of visible posts usr = r"thebeatles" url = "".join([r"https://www.tiktok.com/@", usr]) n = 72 # The most current posts count def get_visible_post_count(page): # Get visible post count post_urls = page.locator("//div[@data-e2e='user-post-item']//a") post_urls_count = len(post_urls.all()) return post_urls_count def press_end(url_n): if not url_n: page.keyboard.press("End") with sync_playwright() as playwright: browser = playwright.chromium.launch() context = browser.new_context() page = context.new_page() helper.logging_network_events(page) # Logging # Abort requests helper.abort_image(page) helper.abort_image_by_ext(page) helper.abort_font_by_ext(page) # Get page page.goto(url, timeout=0) # If nth post element is not visible keep pressing key "End" url_n = page.locator("//div[@data-e2e='user-post-item']//a").nth(n-1).is_visible(timeout=0) for i in range(n): press_end(url_n) # Post count post_count = get_visible_post_count(page) print(post_count) # Save html usr_rel_path = "".join(["html/", usr, "/"]) path_exists = os.path.exists(usr_rel_path) if not path_exists: os.makedirs(usr_rel_path) usr_folder = "".join([os.getcwd(), "/", usr_rel_path]) file_name = "".join([usr, ".html"]) file_dir = "".join([usr_folder, file_name]) html_source = page.content() with open(file_dir, "w") as file: file.write(html_source) file.close() browser.close()
helper.py
import re def logging_network_events(page): # Logging network events page.on("request", lambda request: print( ">>", request.method, request.url, request.resource_type)) page.on("response", lambda response: print( "<<", response.status, response.url)) def abort_image(page): page.route("**/*", lambda route: route.abort() if route.request.resource_type == "image" else route.continue_()) def abort_image_by_ext(page): page.route(re.compile("jpeg|jpg|png|tiff|gif"), lambda route: route.abort()) def abort_font_by_ext(page): page.route(re.compile("woff2|woff|otf"), lambda route: route.abort())
问题分析与解决方法
原代码的核心问题是仅在页面加载完成后检查一次第72个帖子是否可见,之后固定循环按End键,没有考虑TikTok的懒加载机制——每次滚动到底部后,需要等待新帖子加载完成,才能继续滚动。此外,End键一次滚动幅度过大,可能导致页面来不及加载新内容。
修改后的滚动逻辑需要:
- 循环检查当前可见帖子数量,直到达到或超过目标值
- 每次滚动到底部后,等待新帖子加载
- 用更可控的滚动方式触发懒加载
修改后的代码
替换原代码中滚动部分的逻辑,修改后的test_fetch_raw_source.py关键部分如下:
# ... 其他代码保持不变 ... # Get page page.goto(url, timeout=0) # 滚动到目标数量帖子可见的逻辑 target_count = n current_count = get_visible_post_count(page) while current_count < target_count: # 滚动到页面底部,触发懒加载 page.evaluate("window.scrollTo(0, document.body.scrollHeight);") # 等待新帖子加载,可根据实际情况调整等待时间 page.wait_for_timeout(1000) # 更新当前可见帖子数量 current_count = get_visible_post_count(page) print(f"当前可见帖子数: {current_count}") # ... 后续保存HTML代码保持不变 ...
额外优化建议
- 替换固定等待为更精准的元素等待:记录之前的帖子数量,等待数量增加后再继续,避免无效等待。示例:
previous_count = current_count page.evaluate("window.scrollTo(0, document.body.scrollHeight);") # 等待帖子数量增加,超时时间设为5秒 page.wait_for_function( f"document.querySelectorAll('div[data-e2e=\"user-post-item\"] a').length > {previous_count}", timeout=5000 ) current_count = get_visible_post_count(page) - 移除重复的图片请求中止逻辑:原代码中
abort_image和abort_image_by_ext功能重复,保留一个即可,减少不必要的路由处理。
内容的提问来源于stack exchange,提问作者Andrea
相关产品推荐
相关产品推荐

