You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Threading爬取YellowPages时,locations列表仅保留最后一个州数据

问题:多线程爬取Yellow Pages仅获取最后一个州的数据

我尝试从yellowpages.com爬取数据,该网站通过URL https://www.yellowpages.com/state-<state-abbreviation>?page=<letter> 展示某州内以指定字母开头的城市列表,例如纽约州以字母'c'开头的城市对应URL为https://www.yellowpages.com/state-ny?page=c。

需求是将所有“城市,州”组合存入locations变量并写入文件,因单线程效率低改用多线程实现。运行后日志显示完成全部1300个页面(50个州×26个字母)的请求,但文件中仅写入最后一个州怀俄明州(Wyoming)的A-Z城市数据,其他州数据缺失。

我的代码:

def get_session():
    if not hasattr(thread_local, 'session'):
        thread_local.session = requests.Session()
    return thread_local.session

def download_site(url):
    """ Make request to url and scrape data using bs4"""
    session = get_session()
    with session.get(url) as response:
        logging.info(f"Read {len(response.content)} from {url}")
        scrape_data(response)

def download_all_sites(urls):
    """ call download_site() on list of urls"""
    with concurrent.futures.ThreadPoolExecutor(max_workers = 50) as executor:
        executor.map(download_site, urls)


def scrape_data(response):
    """uses bs4 to get city, state combo from yellowpages html and appends to global locations list"""
    soup = BeautifulSoup(response.text, 'html.parser')
    ul_elements = soup.find_all('ul')
    for ul_element in ul_elements:
        anchor_elements = ul_element.find_all('a')
        for element in anchor_elements:
            locations.append(element.text + ',' + state_abbrieviated)

if __name__ == '__main__':
    logging.basicConfig(level=logging.INFO)

    urls = [] # will hold yellowpages urls
    locations = [] # will hold scraped 'city, state' combinations,  modified by scrape_data() function 

    states = {
        'AK': 'Alaska',
        'AL': 'Alabama',
        'AR': 'Arkansas',
        'AZ': 'Arizona',
        'CA': 'California',
        'CO': 'Colorado',
        'CT': 'Connecticut',
        'DC': 'District of Columbia',
        'DE': 'Delaware',
        'FL': 'Florida',
        'GA': 'Georgia',
        'HI': 'Hawaii',
        'IA': 'Iowa',
        'ID': 'Idaho',
        'IL': 'Illinois',
        'IN': 'Indiana',
        'KS': 'Kansas',
        'KY': 'Kentucky',
        'LA': 'Louisiana',
        'MA': 'Massachusetts',
        'MD': 'Maryland',
        'ME': 'Maine',
        'MI': 'Michigan',
        'MN': 'Minnesota',
        'MO': 'Missouri',
        'MS': 'Mississippi',
        'MT': 'Montana',
        'NC': 'North Carolina',
        'ND': 'North Dakota',
        'NE': 'Nebraska',
        'NH': 'New Hampshire',
        'NJ': 'New Jersey',
        'NM': 'New Mexico',
        'NV': 'Nevada',
        'NY': 'New York',
        'OH': 'Ohio',
        'OK': 'Oklahoma',
        'OR': 'Oregon',
        'PA': 'Pennsylvania',
        'RI': 'Rhode Island',
        'SC': 'South Carolina',
        'SD': 'South Dakota',
        'TN': 'Tennessee',
        'TX': 'Texas',
        'UT': 'Utah',
        'VA': 'Virginia',
        'VT': 'Vermont',
        'WA': 'Washington',
        'WI': 'Wisconsin',
        'WV': 'West Virginia',
        'WY': 'Wyoming'
    }
    letters = ['a','b','c','d','e','f','g','h','i','j','k','l','m','n','o',
               'p','q','r','s','t','u','v','w','x','y','z']

    # build list of urls that need to be scrape
    for state_abbrieviated ,state_full in states.items():
        for letter in letters:
            url = f'https://www.yellowpages.com/state-{state_abbrieviated}?page={letter}'
            urls.append(url)

    # scrape data
     download_all_sites(urls)
     logging.info(f"\tSent/Retrieved {len(urls)} requests/responses in {duration} seconds")

     # write data to file
     with open('locations.txt','w') as file:
     for location in locations:
        file.write(location + '\n')

问题根源

  • 全局变量竞态条件:state_abbrieviated是全局变量,构建URL的循环会快速遍历所有州,最终这个变量会被固定为最后一个州的缩写(WY)。多线程并行执行时,所有线程调用scrape_data都会使用这个最终值,导致所有城市错误关联到怀俄明州。
  • 代码缩进错误:原代码中download_all_sites调用、日志输出和文件写入部分存在缩进错误,会导致运行异常(不过你提到日志显示完成所有请求,可能是粘贴时的格式问题)。

修复方案

方案1:绑定URL与州缩写,传递给下载函数

让每个请求携带对应的州缩写,避免依赖全局变量:

def download_site(url, state_abbr):
    """ Make request to url and scrape data using bs4"""
    session = get_session()
    with session.get(url) as response:
        logging.info(f"Read {len(response.content)} from {url}")
        scrape_data(response, state_abbr)

def download_all_sites(url_state_pairs):
    """ call download_site() on list of (url, state_abbr) pairs"""
    with concurrent.futures.ThreadPoolExecutor(max_workers=50) as executor:
        executor.map(lambda x: download_site(*x), url_state_pairs)

def scrape_data(response, state_abbr):
    """uses bs4 to get city, state combo from yellowpages html and appends to global locations list"""
    soup = BeautifulSoup(response.text, 'html.parser')
    ul_elements = soup.find_all('ul')
    for ul_element in ul_elements:
        anchor_elements = ul_element.find_all('a')
        for element in anchor_elements:
            locations.append(f"{element.text},{state_abbr}")

# 构建URL和州缩写的配对列表
url_state_pairs = []
for state_abbrieviated ,state_full in states.items():
    for letter in letters:
        url = f'https://www.yellowpages.com/state-{state_abbrieviated}?page={letter}'
        url_state_pairs.append((url, state_abbrieviated))

# 调用修改后的下载函数
download_all_sites(url_state_pairs)

方案2:从URL中解析州缩写

如果不想修改函数参数,可以直接从响应的URL中提取州缩写:

def scrape_data(response):
    """uses bs4 to get city, state combo from yellowpages html and appends to global locations list"""
    # 从URL中提取州缩写
    state_abbr = response.url.split('state-')[1].split('?')[0]
    soup = BeautifulSoup(response.text, 'html.parser')
    ul_elements = soup.find_all('ul')
    for ul_element in ul_elements:
        anchor_elements = ul_element.find_all('a')
        for element in anchor_elements:
            locations.append(f"{element.text},{state_abbr}")

额外修复:修正代码缩进

确保download_all_sites调用、日志输出和文件写入部分的缩进与if __name__ == '__main__':下的代码对齐:

# 修正后的主函数部分
if __name__ == '__main__':
    logging.basicConfig(level=logging.INFO)

    urls = [] 
    locations = []
    # ... 省略states和letters定义 ...

    # build list of urls that need to be scrape
    url_state_pairs = []
    for state_abbrieviated ,state_full in states.items():
        for letter in letters:
            url = f'https://www.yellowpages.com/state-{state_abbrieviated}?page={letter}'
            url_state_pairs.append((url, state_abbrieviated))

    # scrape data
    import time
    start_time = time.time()
    download_all_sites(url_state_pairs)
    duration = time.time() - start_time
    logging.info(f"\tSent/Retrieved {len(url_state_pairs)} requests/responses in {duration:.2f} seconds")

    # write data to file
    with open('locations.txt','w') as file:
        for location in locations:
            file.write(location + '\n')

内容的提问来源于stack exchange,提问作者Justin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 02:45:36