/
Deep Learning Fundamentals
Save to my account
Sign up
Deep Learning Fundamentals
Deep Learning Fundamentals
Study
1
Question
What is the fundamental purpose of convolutional neural networks regarding data types?
0:40
Answer
Convolutional neural networks are specifically designed for data with spatial structure, such as images.
2
Question
How does the principle of sparse connectivity operate within a convolutional layer?
0:40
Answer
Instead of connecting every neuron to every input, each neuron only processes a local region of the data.
3
Question
What does the mechanism of parameter sharing entail in the context of CNN filters?
0:40
Answer
The same filter is applied across every spatial position, allowing feature detection to occur globally.
4
Question
What is the core objective of the Maximum Likelihood Estimation principle in training?
1:30
Answer
The objective is to find the specific parameters that make the observed data most probable.
5
Question
Why is Mean Squared Error mathematically justified as a loss function for regression?
1:30
Answer
It is a direct consequence of assuming that target data follows a Gaussian distribution centered at the network prediction.
6
Question
How does the stacking of convolutional layers contribute to feature learning?
0:40
Answer
It builds hierarchical representations that transition from simple edges to complex objects as data moves deeper.
7
Question
What is the primary function of pooling layers in a neural network?
0:40
Answer
Pooling layers reduce spatial size and add robustness to small translations within the data.
8
Question
Which loss function corresponds to the categorical distribution used in classification tasks?
1:30
Answer
Cross-entropy is the loss function derived from the statistical assumption of a categorical distribution.
9
Question
What is the definition of overfitting in deep learning models?
2:23
Answer
Overfitting occurs when a model memorizes the training set but fails to generalize to new data.
10
Question
How do L1 and L2 regularization techniques prevent the model from overfitting?
2:23
Answer
They add a penalty for large weights to the loss function, which encourages the model to stay simple.
11
Question
What is the mechanism by which dropout regularization functions during training?
2:23
Answer
It randomly turns off neurons during training to force the network to learn redundant representations.
12
Question
What are the primary components involved in the calculation of a single perceptron?
3:09
Answer
A perceptron multiplies inputs by weights, adds a bias, and applies a nonlinear activation function.
13
Question
Why are nonlinearities essential between the layers of a deep network?
3:09
Answer
Without nonlinearities, the entire network collapses into a single, ineffective linear transformation.
14
Question
Name three common types of nonlinear activation functions used in neural networks.
3:09
Answer
ReLU, sigmoid, and tanh are commonly used nonlinear activation functions.
15
Question
What is the negative consequence of initializing all network weights to zero?
3:09
Answer
Every neuron receives the same gradient and learns identical features, reducing the layer's capacity to a single neuron.
16
Question
What does the Universal Approximation Theorem state about neural networks?
3:09
Answer
It states that a single hidden layer can approximate any continuous function, though it may require exponentially many neurons.
17
Question
What describes the sequence of events during the forward pass of training?
4:25
Answer
Input flows through the layers to produce a prediction and calculate an associated cost.
18
Question
What is the fundamental role of the backpropagation algorithm?
4:25
Answer
Backpropagation is an efficient method for computing the gradient of the cost with respect to every parameter using the chain rule.
19
Question
How does dynamic programming enhance the efficiency of backpropagation?
4:25
Answer
It avoids recomputing the same derivatives by storing intermediate results, which is essential for deep networks.
20
Question
In stochastic gradient descent, what does the learning rate determine?
4:25
Answer
The learning rate controls the magnitude of the step taken in the direction that reduces the cost.
21
Question
How does the RMSProp optimizer differ from standard stochastic gradient descent?
4:25
Answer
RMSProp adapts the step size for each parameter, providing larger steps for small gradients and smaller steps for large ones.
22
Question
What characterizes hierarchical representations in stacked convolutional layers?
0:40
Answer
Early layers detect simple features like edges, while deeper layers represent complex structures like textures and objects.
23
Question
What is the primary objective of regularization techniques in deep learning?
2:23
Answer
The goal is to reduce generalization error and improve performance on new data, even if training error increases.
24
Question
Why is the chain rule from calculus critical for the backpropagation process?
4:25
Answer
It allows the network to propagate derivatives from the cost function backward through each composition of operations to individual weights.
25
Question
How does weight initialization impact the training of a wide neural network?
3:09
Answer
Correct initialization ensures neurons do not learn redundant information by breaking the symmetry of their initial states.
26
Question
What is the role of the bias term within a standard perceptron calculation?
3:09
Answer
The bias is added to the weighted sum of inputs to allow the model to shift the activation function for better fitting.
27
Question
Why are deeper networks generally preferred over shallow networks despite the Universal Approximation Theorem?
3:09
Answer
Deeper networks represent complex functions much more efficiently, requiring fewer total neurons than shallow architectures.
28
Question
How does the training process move from gradient computation to parameter updating?
4:25
Answer
Once gradients are computed via backpropagation, an optimizer takes a step in the direction that reduces the total cost.
29
Question
What is meant by 'translation robustness' in the context of pooling layers?
0:40
Answer
It means the network can still identify a feature even if its position in the image changes slightly.
30
Question
In the Gaussian noise assumption, what does the network's prediction represent?
1:30
Answer
The prediction is treated as the mean of a Gaussian distribution for the target values given the input.