You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从文本文件中提取纯数字并去除千分位分隔符

提取文本中的纯数字2999解决方案

问题场景

文本文件内容示例:

A customer has purchased your 33%-off item for 2,999 dollars

需求:仅提取其中的纯数字2999,已尝试以下代码但不确定后续操作:

def read_file():
    file = open("some.txt", "r")
    numbers = []
    for num in file:
        print(num.split("A customer has purchased your"))

read_file()

解决方案

方法一:正则表达式(灵活通用)

正则能直接匹配数字,还能轻松处理逗号这类分隔符:

import re

def read_file():
    with open("some.txt", "r") as file:
        numbers = []
        for line in file:
            # 先去掉逗号,再匹配所有连续数字
            all_digits = re.findall(r'\d+', line.replace(',', ''))
            # 从匹配结果里挑出目标的4位数字2999
            for digit_str in all_digits:
                if len(digit_str) == 4:
                    numbers.append(digit_str)
                    print(digit_str)
    return numbers

read_file()

方法二:固定格式字符串分割(适合文本结构不变的情况)

如果文本格式一直是...for X dollars这种结构,直接分割定位更简单:

def read_file():
    numbers = []
    with open("some.txt", "r") as file:
        for line in file:
            # 按"for "拆分,取后面的部分
            after_for = line.split("for ")[1]
            # 再按" dollars"拆分,取前面的价格字符串,去掉逗号
            pure_price = after_for.split(" dollars")[0].replace(',', '')
            numbers.append(pure_price)
            print(pure_price)
    return numbers

read_file()

原代码问题说明

你原来的代码里split的参数没加引号(已经帮你补上),但这种分割方式只能把文本拆成两部分,没法直接定位到价格,逻辑上走不通,建议换成上面两种方法。

内容的提问来源于stack exchange,提问作者ok cool

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 07:40:26