Python实现神经网络:Sigmoid函数及相关问题咨询
Great job debugging that weight initialization issue—switching to numpy.random.normal() was a smart fix! Let's break down your questions one by one to help you keep moving forward:
1. Is sigmoid a good activation function?
Sigmoid was a staple in early neural networks, especially for binary classification, because it neatly squashes values into a 0-1 range that’s easy to interpret as probabilities. But it has some significant downsides that make it less ideal for modern deep networks:
- Vanishing gradients: When inputs are very large or small, the derivative of sigmoid gets extremely close to 0. This means during backpropagation, the gradients barely flow to earlier layers—your network essentially stops learning!
- Non-zero-centered outputs: Sigmoid always outputs positive values, which can bias weight updates in one direction, slowing down training.
That said, sigmoid still has its place: it’s perfect for the output layer of binary classification models where you need a probability-like output. For hidden layers, most folks now use ReLU (or variants like Leaky ReLU, GELU) since they avoid vanishing gradients and are faster to compute.
2. What if node outputs are still too large, making sigmoid return 1?
First, fine-tune your weight initialization: Even with numpy.random.normal(), if your standard deviation is too high, you’ll still get oversized weighted sums. Try scaling weights based on the number of inputs—for example, use np.random.normal(0, 1/np.sqrt(input_size), size=...). This keeps activations in a reasonable range by accounting for how many inputs each node receives.
Second, add batch normalization to your network. This layer normalizes inputs to each layer to have a mean of 0 and variance of 1, which prevents activations from blowing up. It’s a huge help for training stable networks and avoids sigmoid saturation at 1 or 0.
Third, if you’re stuck with sigmoid, clip the weighted sums before applying the activation:
weighted_sum = np.clip(weighted_sum, -50, 50) # Sigmoid(-50) ≈ 0, sigmoid(50) ≈ 1, but avoids overflow activation = 1 / (1 + np.exp(-weighted_sum))
Clipping at ±50 is safe because sigmoid is practically 0 or 1 beyond that, but it stops numpy from returning an exact 1.0 due to floating-point limits.
3. How to avoid Python "rounding" very close floats to integers?
Actually, Python isn’t actively rounding these values—it’s a limitation of floating-point precision. When a value is so close to 1.0 that it can’t be represented as a distinct float, numpy returns 1.0. Here’s how to work around it:
- Use higher-precision floats: Numpy supports
np.float128(if your system allows) which has more decimal places, so it can represent values slightly less than 1.0 that would get lost infloat64. - Use a numerically stable sigmoid implementation: Replace your naive sigmoid with
scipy.special.expit(weighted_sum). This function handles large x values without overflow and returns the closest possible float to the true sigmoid value, instead of defaulting to 1.0. - Clip your inputs: As mentioned earlier, keeping weighted sums within ±30-50 ensures sigmoid outputs stay distinguishable from 1.0 in standard float formats.
4. Advice for a neural network beginner
- Books: You already started with Make Your Own Neural Network—excellent for building intuition! Next, try Neural Networks and Deep Learning by Michael Nielsen for a deeper, math-focused dive that’s still accessible. For practical, code-heavy learning, Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow by Aurélien Géron is a must—it walks you through real projects and teaches you when to use which techniques.
- Build small projects: Don’t just read—get your hands dirty! Start with MNIST digit classification (the "hello world" of neural networks) using your custom class, then experiment with frameworks like Keras to see how they simplify things. Mess with activation functions, weight initializations, and network sizes to learn what works.
- Learn the core math: Skipping calculus and linear algebra might seem tempting, but understanding backpropagation (chain rule, gradients) will help you debug issues like the one you faced. Even a basic grasp of derivatives and matrix multiplication goes a long way.
- Join communities: Hang out in places like Stack Overflow, Reddit’s r/MachineLearning, or beginner-focused Discord groups. Asking questions (like you did here!) and seeing how others solve problems is one of the fastest ways to learn.
- Start shallow: Don’t jump straight into deep learning. Master single-layer and small multi-layer networks first, then move to CNNs or RNNs once you’re comfortable with the basics.
内容的提问来源于stack exchange,提问作者user9363390

