You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Python Web代理的DNSLookupFailedError问题

Hey there! Let's work through this proxy resource matching issue you're facing. It sounds like your transform_content.py is missing patterns for certain resource URLs, causing errors when accessing specific sites. Here's a breakdown of steps to fix this, plus alternative library options if you want to simplify things:

1. First, Identify the Missing Resource Patterns

Before tweaking code, figure out exactly which URLs aren't being caught:

  • Open your browser's DevTools > Network tab when accessing the problematic site.
  • Look for requests that return 404 errors, or that load directly from the original domain instead of your proxy.
  • Note the URL formats: are they protocol-relative (like //example.com/script.js), relative paths with ../, special resource types (.woff2 fonts, .webp images), or hidden in attributes like data-src/srcset?
2. Audit Your Current URL Matching Logic in transform_content.py

Chances are your existing regex or string matching isn't covering edge cases. Common gaps include:

  • Missing support for single-quoted attributes (e.g., src='image.png' instead of src="image.png").
  • Ignoring CSS background: url('/bg.jpg') style declarations.
  • Not handling srcset attributes for responsive images.
  • Skipping protocol-relative URLs or root-relative paths.

For example, if your current code uses a regex like r'src="([^"]+)"', it won't catch single-quoted values or unquoted attributes. You can expand it to cover more cases:

# Match src/href with single quotes, double quotes, or no quotes
content = re.sub(r'(src|href)=["\']?([^"\'>\s]+)["\']?', lambda m: f'{m.group(1)}="{proxy_url(m.group(2))}"', content)
3. Add Coverage for Edge Case Resource Types

Here are specific patterns to add to your transformation logic:

  • CSS background URLs:
    # Replace url() values in inline styles and style tags
    content = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2))}")', content)
    
  • Responsive image srcset:
    # Split srcset values and proxy each URL
    def proxy_srcset(srcset):
        return ', '.join([f'{proxy_url(url)} {size}' if size else proxy_url(url) for url, size in (s.strip().split(' ', 1) for s in srcset.split(','))])
    
    content = re.sub(r'srcset=["\']?([^"\'>]+)["\']?', lambda m: f'srcset="{proxy_srcset(m.group(1))}"', content)
    
  • Data attributes like data-src:
    content = re.sub(r'data-src=["\']?([^"\'>\s]+)["\']?', lambda m: f'data-src="{proxy_url(m.group(2))}"', content)
    
4. Alternative Libraries to Simplify URL Transformation

If maintaining regex patterns feels brittle, use dedicated HTML parsing libraries to reliably traverse and modify content:

Option 1: BeautifulSoup

This library makes it easy to iterate over all HTML tags and attributes without regex headaches:

from bs4 import BeautifulSoup

def transform_content(content, base_url, proxy_prefix):
    soup = BeautifulSoup(content, 'html.parser')
    
    # Process standard resource tags
    for tag in soup.find_all(['img', 'script', 'link', 'a', 'source']):
        for attr in ['src', 'href', 'srcset']:
            if tag.get(attr):
                tag[attr] = proxy_url(tag[attr], base_url, proxy_prefix)
    
    # Process inline styles and style tags
    for style_tag in soup.find_all('style'):
        if style_tag.string:
            style_tag.string = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2), base_url, proxy_prefix)}")', style_tag.string)
    
    for tag in soup.find_all(style=True):
        tag['style'] = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2), base_url, proxy_prefix)}")', tag['style'])
    
    return str(soup)

Option 2: lxml

For faster parsing of large content, lxml is a great alternative (works similarly to BeautifulSoup but with better performance):

from lxml import html

def transform_content(content, base_url, proxy_prefix):
    tree = html.fromstring(content)
    
    # Process attributes
    for attr in ['src', 'href', 'srcset']:
        for element in tree.xpath(f'//*[@{attr}]'):
            original_url = element.get(attr)
            element.set(attr, proxy_url(original_url, base_url, proxy_prefix))
    
    # Process CSS styles
    for style_element in tree.xpath('//style'):
        if style_element.text:
            style_element.text = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2), base_url, proxy_prefix)}")', style_element.text)
    
    return html.tostring(tree, encoding='unicode')
5. Use Your Test Suite to Validate Fixes

Since you've added a test suite, create test cases that cover the problematic URL patterns you identified. For example:

  • Test protocol-relative URLs (//example.com/style.css)
  • Test relative paths with ../ (../assets/logo.png)
  • Test CSS background URLs and srcset attributes

Run the tests after each change to ensure your fixes don't break existing functionality.

If you can share specific error logs or examples of URLs that aren't being proxied, we can narrow down the exact pattern to add!

内容的提问来源于stack exchange,提问作者T145

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:01:15