You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python正则提取的代理地址与对应时间记录配对组合为指定格式

How to Pair Proxies with Corresponding Time Records

The issue with your current code is that it yields all proxy addresses first, then all time records—so they’re output as two separate groups instead of paired one-to-one. To fix this, we can collect both sets of data into lists first, then combine them by zipping the lists together. Here’s how to adjust your code:

import re

def tst():
    text = ''' <script> '''  # Your original HTML text here
    proxies = []
    timestamps = []
    
    # Collect proxy addresses
    if proxy_matches := re.findall(r"(?:<td\s[^>]*?><font\sclass=spy14>(.*?)<script.*?\\"\+(.*?)</script)", text):
        for proxy, port in proxy_matches:
            proxies.append(f"{proxy}:{''.join(port)}")
    
    # Collect time records
    if time_matches := re.findall(r"<td colspan=1><font class=spy1><font class=spy14>(.*?)</font> (\d+[:]\d+) <font class=spy5>([(]\d+ \w+ \w+[)])", text):
        for date, time, taken in time_matches:
            timestamps.append(f"{date} {' '.join([time, taken])}")
    
    # Pair and yield combined entries
    for proxy, timestamp in zip(proxies, timestamps):
        yield f"{proxy} - {timestamp}"

for entry in tst():
    print(entry)

Key Changes:

  1. Collect Data into Lists: Instead of yielding each proxy/time immediately, we store them in proxies and timestamps lists.
  2. Zip and Combine: Using zip(proxies, timestamps) pairs each proxy with the corresponding time record (assuming the number of proxies and times are equal, which they are in your example).
  3. Yield Paired Strings: We generate the combined format you want directly in the final loop.

Expected Output:

51.155.10.0:8000 - 27-oct-2022 11:05 (49 mins ago)
178.128.96.80:7497 - 27-oct-2022 11:04 (50 mins ago)
98.162.96.41:4145 - 27-oct-2022 11:03 (51 mins ago)

This approach keeps your existing regex logic intact while ensuring the data is paired correctly. If you ever encounter a mismatch in the number of proxies and times, you could add a check (like using itertools.zip_longest to handle missing entries), but this should work perfectly for your current use case.

内容的提问来源于stack exchange,提问作者xnoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 13:29:08