如何用Python遍历文本文件列表并批量使用Linkfinder处理多URL
问题解答
一、Python遍历文本文件中的列表
分两种常见场景处理:
场景1:文本文件每行存储一个元素(比如你的JS URL)
直接按行读取,跳过空行即可:
# 假设文件名为js_urls.txt with open('js_urls.txt', 'r', encoding='utf-8') as f: for line in f: item = line.strip() # 去除换行符和前后空白 if item: # 跳过空行 print(item) # 这里替换成你需要的处理逻辑
场景2:文本文件内容是Python列表格式(如["url1", "url2"])
用ast.literal_eval安全解析列表:
import ast with open('js_urls_list.txt', 'r', encoding='utf-8') as f: content = f.read().strip() item_list = ast.literal_eval(content) for item in item_list: print(item) # 执行你的处理逻辑
二、用Linkfinder批量处理多个JS URL
可以通过脚本循环调用Linkfinder,避免手动逐个执行:
方法1:Python脚本批量调用
import subprocess with open('js_urls.txt', 'r', encoding='utf-8') as f: for idx, line in enumerate(f, 1): url = line.strip() if not url: continue # 每个URL生成独立结果文件,防止覆盖 output_file = f'result_{idx}.html' # 执行Linkfinder命令 cmd = f'python linkfinder.py -i "{url}" -o "{output_file}"' subprocess.run(cmd, shell=True) print(f'处理完成:{url} -> 结果文件:{output_file}')
方法2:Bash脚本(Linux/macOS适用)
创建batch_linkfinder.sh文件:
#!/bin/bash idx=1 while read -r url; do if [ -n "$url" ]; then python linkfinder.py -i "$url" -o "result_$idx.html" echo "处理完成:$url -> result_$idx.html" idx=$((idx+1)) fi done < js_urls.txt
添加执行权限并运行:
chmod +x batch_linkfinder.sh ./batch_linkfinder.sh
额外处理:如果文本文件含HTML标签(如你的示例<a>标签)
先提取纯URL再处理:
from bs4 import BeautifulSoup with open('urls_with_html.txt', 'r', encoding='utf-8') as f: content = f.read() soup = BeautifulSoup(content, 'html.parser') urls = [a['href'] for a in soup.find_all('a')] # 遍历urls列表执行批量处理逻辑即可
内容的提问来源于stack exchange,提问作者tony michael
相关产品推荐
相关产品推荐

